Summary
Recently I have been looking into ways to test my Apache Beam pipelines at work. Most common use cases of Beam generally involve either batch reading data from GCS and writing to analytical platforms such as Big Query or stream reading data from Pubsub and writing to perhaps Bigtable.
A pipeline consists of transforms and its generally easy to test them in isolation as an independent unit test per stage, however I am personally a big fan of “end-to-end” testing or “Integration testing” and this is where things can sometimes get tricky.