Data Engineer
Fitch Ratings
Owned ETL pipelines using AWS Glue and Python, collaborating with stakeholders to load data warehouses efficiently and improve data availability by 40%. Designed Apache Airflow workflows in Python to orchestrate complex data pipelines, reducing manual errors and increasing processing reliability across systems. Architected PySpark jobs leveraging Spark to accelerate large-scale data transformations, cutting ingestion times by 35% for compliance-driven data workflows. Standardized transformation layers with dbt and defined consistent data contracts, ensuring modular pipeline scalability and reliable downstream analytics. Built Python and SQL-based data quality frameworks that tested pipelines rigorously, achieving 99.5% accuracy and reducing data reconciliation errors by 25%. Optimized complex SQL queries in Amazon Redshift, improving query execution times by 35% during peak processing windows in a multi-terabyte data environment. Provisioned cloud infrastructure using Terraform on AWS, automating environment setup to enhance deployment consistency and scalability across data platforms. Integrated GitHub Actions with AWS CodePipeline to automate pipeline deployment and transformation validation, reducing release errors and accelerating delivery cycles. Designed RESTful APIs in Python for real-time data ingestion, increasing update frequency of environmental, social, and governance metrics by 50%. Led ETL pipeline redesign with a scalability focus, proactively reducing processing latency by 28% and improving system reliability for high-volume workloads.