Start here →

National Grocer Converts 5,100 SAS Jobs and 4.4 Million Lines to Open Apache PySpark on EMR and Glue

SAS2PY Case Study • March 2026 • Grocery & CPG Retail

Executive Summary

A US grocery retailer with 2,400+ stores already ran a Hadoop leftover and a growing AWS analytics footprint (EMR, Glue, S3, MWAA). What they did not want was a second platform tax. SAS Grid plus DI Studio was still the system of record for replenishment, pricing, loyalty, and vendor scorecards — 5,100 scheduled jobs, 4.4 million lines of SAS, 2,100 autocall macros, and a 36-node Grid that was out of hardware support. MigryX parsed the estate and emitted Apache PySpark 3.5 (DataFrame API + Spark SQL) that runs on EMR Serverless for the heavy DAGs and AWS Glue for the short, high-cardinality jobs. Sixteen months later the Grid is powered off. Projected three-year savings: $7.4 million, almost all of it SAS license + Grid ops, not “cloud magic.”

Client Overview

Merchandising, replenishment, and loyalty analytics grew up in SAS because that is what the 2008 data-warehouse program standardized on. Enterprise Guide authored most analyst jobs. DI Studio owned the 1,180 nightly vendor and POS ingest flows. A smaller set of stored processes fed store-ops dashboards. The company had already moved raw POS and loyalty events to S3; SAS was still the place those events became decisions.

Platform engineering had evaluated Databricks and Snowflake. Both were rejected for this estate: the firm already paid for EMR capacity commitments, the security team had finished a Lake Formation rollout, and merchandising leadership would not accept a second catalog and a second IAM story. The requirement was explicit — open PySpark, their S3, their Glue catalog, their Airflow.

Business Challenge

The MigryX Approach

Inventory first: hash every program, collapse parameterized Control-M clones into one logical job with a parameter file, and build the DI metadata graph. That reduced the conversion surface from 5,100 entries to 3,640 programs + 1,180 flows — still large, but no longer padded.

The AST path is the same SAS parser used on Databricks programs. The emitter is not. Output is plain PySpark: spark.read / DataFrame transforms / spark.sql for PROC SQL, written as importable modules with a thin runner. No Databricks Workflow YAML, no Unity Catalog names, no dbutils. Table identifiers are Glue database.table. Storage is S3 prefixes that match the old SAS library names so merchandisers can find “PRD.LOYALTY” without a decoder ring.

Runtime placement is a classifier, not a preference:

MWAA DAGs were generated from Control-M + DI predecessors, then patched where the scheduler and the metadata disagreed (they did, on 94 flows). That patch list is part of the deliverable; pretending the graph was clean would have shipped a lying DAG.

Target Architecture

SAS Grid → MigryX → Apache Spark on EMR + Glue (no Databricks control plane)

On-premisesOn-premises
SAS GridSAS Grid36 nodes, Control-M
DI StudioDI Studio1,180 flows
EG + STPEG + STPAnalyst jobs
MigryXMigryX
PySpark 3.5PySpark 3.5No dbutils
Runtime hintRuntime hintEMR vs Glue
AWS us-east-1AWS us-east-1
Amazon S3Amazon S3LIBNAME prefixes
EMR ServerlessEMR ServerlessHeavy shuffle
AWS GlueAWS GlueShort fan-out
MWAAMWAAControl-M graph
Lake FormationLake FormationLoyalty PII
CloudTrailCloudTrailJob + table audit

This is Spark-the-engine, not Spark-the-product. If they later put Iceberg under these tables, the PySpark modules do not change — only the catalog writer. That was a design constraint, not a slide.

What we refused to emit

collect() to pandas for “simple” PROCs, rdd.map for DATA steps, and UDFs for SAS formats. Formats became broadcast lookup DataFrames (the 94-format catalog was 11 MB). DATA-step RETAIN became windows. The 71 programs that were genuinely sequential (point= loops over claim-like store incidents) stayed as mapPartitions with an explicit comment and a size cap — they are 71, not a license to write RDD code everywhere.

Estate Inventory and Cutover Waves

Domain Jobs (logical) LOC Runtime Wave
POS / replenishment 1,280 1.1M EMR Serverless 1–2
Vendor / ASN / compliance 1,180 820k Glue (fan-out) 2
Loyalty / identity 740 690k EMR + Lake Formation 3
Pricing / promo 610 540k EMR 3–4
Finance / vendor pay 420 410k EMR, dual-run 45 days 4
Store-ops extracts / STP 590 840k Glue SQL-only where possible 5

4.4M LOC includes macros once, not once per caller. DI XML is counted as generated SAS equivalent, because that is what the Grid actually ran. We do not count Control-M JCL as “code.”

01CollapseParameter clones → one module + JSON params
02ConvertPySpark module + runner, no dbutils
03PlaceHint: EMR / Glue / SQL from Grid profiles
04ParityItem-store key hashes on 28 days of POS
05CutMWAA live, Control-M entry disabled

What actually moved the needle

Results

5,100
Scheduled jobs (3,640 + 1,180 logical)
4.4M
Lines of SAS converted once
3.1X
Replenishment DAG wall-clock
$7.4M
3-year license + Grid gap
16 mo
Inventory through Grid power-off
88%
No human rewrite
"We already knew Spark. We did not need a new logo on the architecture diagram. We needed 5,000 SAS jobs to become modules our on-call could grep. The parser did the MERGE cases we were going to get wrong by hand."

— Director of Data Platform, US grocery retailer

SAS estate, open Spark runtime

EMR, Glue, Dataproc, Cloudera, or plain Spark — same parser, no platform lock-in in the emitted code.

Explore PySpark modernization →

SAS modernization paths

Targets: Databricks Snowflake Google Cloud Azure AWS PySpark Polars Iceberg