Spark DataSet - DSL Operations

About

Domain-specific-language (DSL) functions are defined in the class:

Example:

  • group by,
  • order,
  • plus,….

Management

Pipeline

With a spark session and a dataset of row (ie Spark DataSet - Data Frame)

  • Scala
val people = spark.read.parquet("...")
val department = spark.read.parquet("...")

people.filter("age > 30")
 .join(department, people("deptId") === department("id"))
 .groupBy(department("name"), people("gender"))
 .agg(avg(people("salary")), max(people("age")))
  • Java:
Dataset<Row> people = spark.read().parquet("...");
Dataset<Row> department = spark.read().parquet("...");

people.filter(people.col("age").gt(30))
 .join(department, people.col("deptId").equalTo(department.col("id")))
 .groupBy(department.col("name"), people.col("gender"))
 .agg(avg(people.col("salary")), max(people.col("age")));

Powered by ComboStrap