During my data engineering internship with Explore Data Science Academy, much of my time went into ETL pipelines: extracting data from several sources, transforming it into something useful, and loading it where analysts and models could use it. These are the lessons that stuck.

Start with plain Python

Spark is powerful, but it's overkill for a few hundred megabytes. I start with pandas and only move to PySpark when the data or the processing time genuinely demands it. The structure of the pipeline stays the same either way.

Make every step re-runnable

A pipeline will fail halfway through at some point. If re-running it creates duplicates, you have a bigger problem than the failure. Write each step so that running it twice gives the same result:

(df.write
   .mode("overwrite")
   .partitionBy("event_date")
   .parquet("s3://warehouse/events/"))

Overwriting one date partition at a time means a rerun replaces that day's data instead of appending to it.

Separate raw from clean

Keep an untouched copy of the raw input. When a transformation turns out to be wrong, and one eventually will, you can rebuild the clean tables from the source instead of asking the source system for history it may no longer have.

Test the data

Unit tests catch bugs in code. They don't catch a supplier who starts sending dates in a new format. Add checks at the boundaries:

  • Row counts within an expected range
  • No nulls in key columns
  • Values inside sensible bounds

Fail the run when a check fails. Loading bad data quietly is worse than loading nothing.

Schedule and observe

Use a scheduler, not a cron job on someone's laptop, and record how long each run takes and how many rows it moved. When a run that normally takes ten minutes takes two hours, you want to know before the morning report goes out.