ML Engineer MasterClass (October) | 4 seats left

Spark · Maps and JSON
AmazonAmazon Analytics
Amazon · Work with Nested Data

Maps and JSON

Look up map keys and parse JSON into typed fields.

Step 1 of 6 · Learn

Read map values

attributes maps keys to string values. Both status and customer are strings, so all map values share a type. getItem reads by key rather than position. A missing key returns null; a map does not have the fixed named-field schema of a struct.

Lesson reference: PySpark and Scala

Read map values

attributes maps keys to string values. Both status and customer are strings, so all map values share a type. getItem reads by key rather than position. A missing key returns null; a map does not have the fixed named-field schema of a struct.

PySpark example

from pyspark.sql import Window
from pyspark.sql.functions import (
    col, lit, array, when, explode, explode_outer, posexplode,
    create_map, to_json, struct, from_json, get_json_object, count, row_number
)

map_orders = orders.select(
    "order_id", create_map(lit("status"), col("status"),
    lit("customer"), col("customer_id").cast("string")).alias("attributes")
)

result = map_orders.select(
    col("order_id"), col("attributes").getItem("status").alias("attribute_value")
).orderBy("order_id")
result.show(truncate=False)

Scala example

import org.apache.spark.sql.expressions.Window
import org.apache.spark.sql.functions._

val map_orders = orders.select(
  col("order_id"), map(lit("status"), col("status"),
  lit("customer"), col("customer_id").cast("string")).alias("attributes")
)

val result = map_orders.select(
    col("order_id"), col("attributes").getItem("status").alias("attribute_value")
).orderBy("order_id")
result.show(false)

Parse with a schema

The setup serializes status and item_count to JSON text. from_json uses the supplied schema to return a struct with a string and an integer. Missing fields become null; malformed input needs a deliberate parsing policy. These examples use valid JSON with matching types.

PySpark example

from pyspark.sql import Window
from pyspark.sql.functions import (
    col, lit, array, when, explode, explode_outer, posexplode,
    create_map, to_json, struct, from_json, get_json_object, count, row_number
)

json_orders = orders.select(
    "order_id", to_json(struct("status", "item_count")).alias("payload")
)

result = json_orders.withColumn(
    "parsed", from_json(col("payload"), "status STRING, item_count INT")
).select(col("order_id"), col("parsed.status").alias("status")).orderBy("order_id")
result.show(truncate=False)

Scala example

import org.apache.spark.sql.expressions.Window
import org.apache.spark.sql.functions._

val json_orders = orders.select(
  col("order_id"), to_json(struct("status", "item_count")).alias("payload")
)

val result = json_orders.withColumn(
    "parsed", from_json(col("payload"), "status STRING, item_count INT")
).select(col("order_id"), col("parsed.status").alias("status")).orderBy("order_id")
result.show(false)

Extract a JSON path

get_json_object extracts a path as text. Even when item_count is numeric in the JSON, the extracted value is a string. Cast it when a numeric result is required, or use from_json to type several fields together. All item_count values here are valid integers.

PySpark example

from pyspark.sql import Window
from pyspark.sql.functions import (
    col, lit, array, when, explode, explode_outer, posexplode,
    create_map, to_json, struct, from_json, get_json_object, count, row_number
)

json_orders = orders.select(
    "order_id", to_json(struct("status", "item_count")).alias("payload")
)

result = json_orders.select(
    col("order_id"), get_json_object(col("payload"), "$.item_count").alias("items")
).orderBy("order_id")
result.show(truncate=False)

Scala example

import org.apache.spark.sql.expressions.Window
import org.apache.spark.sql.functions._

val json_orders = orders.select(
  col("order_id"), to_json(struct("status", "item_count")).alias("payload")
)

val result = json_orders.select(
    col("order_id"), get_json_object(col("payload"), "$.item_count").alias("items")
).orderBy("order_id")
result.show(false)
example.pyPySpark
Map / JSONSelect a key or fieldTyped output
Source map_orders6 rows
order_idattributes
1001{"status":"Delivered","customer":"101"}
1002{"status":"Shipped","customer":"102"}
1003{"status":"Cancelled","customer":"101"}
1004{"status":"Delivered","customer":"103"}
1005{"status":"Delivered","customer":"104"}
1006{"status":"Shipped","customer":"105"}
Read map values
Result6 rows
order_idattribute_value
1001Delivered
1002Shipped
1003Cancelled
1004Delivered
1005Delivered
1006Shipped
The key chooses a value inside each map; the number of order rows stays unchanged.

The example is loaded in the editor. Run it as written, then try a small change.

Runs on the Spark backend. First startup may take a moment.

Run your code to see the result.