sai krishna

Data Engineer

IL, United States

#OpenToWork

About

I am a Data Engineer with experience designing and optimizing scalable ETL pipelines using Airflow, dbt, and Spark within AWS environments. Most recently at Fitch Ratings, I have been working since 2022, giving me over 2 years of experience in data engineering roles. My work spans Python, SQL, Scala, and Bash scripting to build and maintain data pipelines. I have built Apache Airflow workflows, PySpark jobs, and dbt transformation layers, while managing AWS Glue, Lambda, S3, Redshift, and Terraform infrastructure. I have also automated deployments using GitHub Actions and AWS CodePipeline and designed RESTful APIs for real-time data ingestion.

Experience

Fitch Ratings

Data Engineer

Fitch Ratings

Apr 2023 – Present

 Owned ETL pipelines using AWS Glue and Python, collaborating with stakeholders to load data warehouses efficiently and improve data availability by 40%.  Designed Apache Airflow workflows in Python to orchestrate complex data pipelines, reducing manual errors and increasing processing reliability across systems.  Architected PySpark jobs leveraging Spark to accelerate large-scale data transformations, cutting ingestion times by 35% for compliance-driven data workflows.  Standardized transformation layers with dbt and defined consistent data contracts, ensuring modular pipeline scalability and reliable downstream analytics.  Built Python and SQL-based data quality frameworks that tested pipelines rigorously, achieving 99.5% accuracy and reducing data reconciliation errors by 25%.  Optimized complex SQL queries in Amazon Redshift, improving query execution times by 35% during peak processing windows in a multi-terabyte data environment.  Provisioned cloud infrastructure using Terraform on AWS, automating environment setup to enhance deployment consistency and scalability across data platforms.  Integrated GitHub Actions with AWS CodePipeline to automate pipeline deployment and transformation validation, reducing release errors and accelerating delivery cycles.  Designed RESTful APIs in Python for real-time data ingestion, increasing update frequency of environmental, social, and governance metrics by 50%.  Led ETL pipeline redesign with a scalability focus, proactively reducing processing latency by 28% and improving system reliability for high-volume workloads.

AWS GlueApache AirflowPySpark

Education

Texas A&M University

Texas A&M University

M.S. · Computer Science

2021 – 2022

Certifications

AWS Cloud Practitioner

AWS

Skills

RESTful APIsGitHubData QualityData LakesData WarehousesDockerTerraformSnowflakePostgreSQLAWS RedshiftAWS S3AWS LambdaAWS GluedbtPySparkApache AirflowETLScalaSQLPython

Languages

English (Professional working proficiency)