SHAHRUKH KHAN
Lead Data Engineer
Kichha, India
#OpenToWork
About
•Adaptable professional with Overall 8 years of experience in Data Engineering, Big Data and Data Warehousing technologies, including 2+ years in a Lead Data Engineer role.
•Extensive hands-on experience with Big Data solutions, including Hadoop, HDFS, MapReduce (MR), Spark, Hive, HBase, Sqoop, Python and PySpark.
•Proficient in working with Spark APIs, including Spark RDD, Dataframe and Dataset APIs.
•Worked with various file formats such as JSON, PARQUET, AVRO, ORC and text files.
•Expertise in Extraction, Transformation, and Loading (ETL) using Sqoop for seamless data migration between RDBMS and HDFS.
•Designed and implemented scalable and reliable data pipelines utilizing Azure services like Azure Data Factory, Azure Databricks, Azure Synapse Analytics, and Azure SQL.
•Utilized AWS technologies including EC2, Lambda, EMR, Athena, Glue, S3, Redshift and IAM Policies.
•Implemented ETL pipelines and data transformations using Snowflake and DBT for modern cloud data warehouse solutions.
•Repository usage – Git hub, Bit bucket
•Familiarity with agile methodologies and a proven track record of working in agile development environments.
•Strong communication and collaboration skills, adept at working in cross-functional teams to deliver data solutions within defined timelines and budgets.
•Excellent team player with in-depth knowledge of development tools and languages.
What I'm looking for
Looking for a full-time remote role as a Lead or Senior Data Engineer. Experienced in building scalable data platforms, cloud-native ETL pipelines, and modern data warehouses using Azure, AWS, Databricks, Apache Spark, PySpark, Python, and SQL. Interested in solving complex data challenges, mentoring teams, and delivering reliable, high-performance data solutions.
Experience
Lead Data Engineer
BrightCanyon Solutions India
Aug 2018 – Present
Project: Insurance Data Analysis Client: Liberty Mutual
01/2024 – Present
Team Size: 6
Environment: Agile Methodology, AWS (IAM, EC2, S3, Athena, Glue, EMR, Redshift), PySpark, Hadoop, HDFS, Hive, Airflow, Sqoop, MySQL, Python, Teradata, SAP HANA, Salesforce, Git, Jira Project Details: The objective of this project is to build a centralized insurance data platform where policy, claims, customer and agent related data from multiple source systems is collected, processed and prepared for analytics, reporting and fraud detection use cases. Data is ingested into AWS S3, transformed using AWS Glue and PySpark on EMR, and loaded into Amazon Redshift for business reporting and analytics.
Roles & Responsibilities:
•Design and implemented data pipelines using AWS services such as S3, Glue, and EMR.
•Loaded and Transformed large sets of Structured, Semi Structured and Unstructured data.
•Developing and maintaining technical documentation for Spark applications and data processing workflows, including data models, ETL processes, and system architecture.
•Experience in working with big data technologies such as Hadoop, Spark, and Kafka, and integrating them with AWS services.
•Proficient in programming languages such as Python for developing Spark applications and big data solutions.
•Developed PySpark ETL jobs using AWS Glue and EMR to ingest, transform and process data from multiple source systems and load curated datasets into Amazon Redshift.
•Designed and maintained Redshift tables, implemented data loading strategies and optimized analytical queries for reporting, insurance analytics and business consumption.
•Was responsible for Optimizing Spark SQL queries that helped in saving Cost to the project.
•Having hands-on experience with Spark Memory Tuning.
•Working with stakeholders to understand business requirements and translating them into technical requirements for big data solutions.
AWS GluePySparkAmazon RedshiftAWS S3EMR
Associate Consultant-DATA ENGINEER
Brightcanyon Solutions
Aug 2018 – Dec 2023
Project: MARKETING ANALYTICS PLATFORM Client: HSBC Bank
01/2021 – 12/2023
Team Size: 4
Environment: Agile methodology, Azure(boards, ADF, ABS, ADB, ASD, Azure Synapse Analytics), Pyspark, HDFS, Hadoop, Hive, Airflow, Sqoop, Python, MySQL, PostgreSQL, GIT, Jira Project Details: The goal of this project was to develop a robust Marketing Analytics Platform for a bank, enabling data-driven marketing strategies and insights to drive customer acquisition, retention, and campaign optimization. As a Data Engineer, I played a key role in designing and implementing the data engineering components of the platform.
Roles & Responsibilities:
•Designing and implementing data processing workflows using Azure Data Factory, including ETL processes and data movement between Azure services and on-premises data sources.
•Developing and optimizing Spark SQL queries to extract insights from data stored in Azure Blob Storage or Azure Data Lake Storage.
•Proficient in data modeling and database design using Azure services such as Azure SQL Database and Azure Cosmos DB.
•Involved in data clean-up by removing duplicated data.
•Involved in Collecting Business Requirements from Business Users, Translate into Technical Design (Data Pipelines and ETL workflows).
•Experience in working with big data technologies such as Hadoop, Spark, and Kafka, and integrating them with Azure services.
•Involved in import data from various RDBMS into Hive Tables which includes Queries using Sqoop.
•Generated and processed complex JSON data after all the transformations for easy storage and access as per client requirements.
•Development of Code & peer review of assigned task and Bug fixing.
Spark SQLSqoopDBTSnowflakeSQL
Education
Kumaun University
Bachelor of arts (B.A)
2016
Skills
Azure Data FactoryAWS LambdaAWS GlueAWS S3AWS IAMAWS EC2SAP HANAPostgreSQLTeradataMySQLDBTSnowflakeAirflowApache SparkHiveSqoopHBaseMapReduceHDFSHadoop
Languages
English (Full professional proficiency)Hindi (Native or bilingual proficiency)