Spark DataSet - DSL Operations

Domain-specific-language (DSL) functions are defined in the class:


  • group by,
  • order,
  • plus,….



With a spark session and a dataset of row (ie Spark DataSet - Data Frame)

  • Scala
val people ="...")
val department ="...")

people.filter("age > 30")
 .join(department, people("deptId") === department("id"))
 .groupBy(department("name"), people("gender"))
 .agg(avg(people("salary")), max(people("age")))
  • Java:
Dataset<Row> people ="...");
Dataset<Row> department ="...");

 .join(department, people.col("deptId").equalTo(department.col("id")))
 .groupBy(department.col("name"), people.col("gender"))
 .agg(avg(people.col("salary")), max(people.col("age")));

