Fictional resume example. Names, employment histories and results are illustrative, not an actual employee record.
Rebecca Shaw
Data Engineer with experience in reliable pipelines, data quality, platforms, and analyst productivity. Practical work includes SQL, Data pipelines, Source-to-target reconciliation, Partition replay.
Experience
Salesforce
San Francisco, California · United States
Data Engineer
Mar 2022 - now
- Built Azure data pipelines that load operational extracts into curated tables, using SQL merge logic with explicit keys and late-arrival rules; tested reprocessing without duplicating records.
- Added SQL reconciliation and quarantine checks to the data pipelines, retaining rejected rows with a reason code; traced an upstream schema change before it reached the reporting layer.
- Added source-to-target row reconciliation and late-arrival handling to a data pipeline; replayed an interrupted partition without duplicating completed records.
- Owned an Airflow ETL recovery path using partitioned SQL fixtures, resolved duplicate loads and completed replay without changing downstream business keys.
- Built Python data-pipeline checks using Spark row reconciliation, reduced late-partition investigation from 45 to 18 minutes and published the Azure rerun guide.
Microsoft
Redmond, Washington · United States
Data Engineer
Jan 2019 - Feb 2022
- Published certified datasets and ownership for 15 recurring business questions. Kept turnaround time within the agreed operating window.
- Investigated a discrepancy between a published metric and its source records, traced the transformation that changed the population, and corrected the calculation with a reproducible query.
- Compared the last successful data refresh with a failed run, separated missing source data from transformation errors, and reran only the affected interval. Kept the original query and corrected result together for review.
Selected project
Data Engineer — independent case study
Project owner
Feb 2024 - Jun 2024
- Implemented incremental models and partition pruning for the 12 largest transforms
- Generated synthetic source records with duplicates, late arrivals and corrected values; wrote assertions for row counts and key uniqueness, and recorded the expected effect of each case on the reported metric.
- Compared the analytical output with a manually calculated reference table, traced differences to a transformation step, and retained a data dictionary and rerun instructions alongside the corrected query.
- Owned the synthetic-data validation using a manually calculated reference, resolved duplicate-key inflation and completed a notebook that reproduces the corrected totals.
Education
University of Washington
Seattle, Washington · United States
B.S. Computer Science
Sep 2013 - Jun 2017
Relevant coursework: Algorithms, operating systems, databases, computer networks
Skills
Role expertise
SQL · Data pipelines · Source-to-target reconciliation · Partition replay
Publications
- Published an independent methods note using reproducible queries, explaining the data grain, excluded records and sensitivity of the result to a changed denominator.


