Kirtesh Thakre

Kirtesh Thakre

Team Lead – AWS Cloud Infrastructure | Cloud Operations | AWS Support | Incident Management | Platform Reliability | Linux | Terraform | Production Support | 10 Years | SRE | Cloud Security | DevOps

Nagpur, India

#OpenToWork

About

Cloud Infrastructure Operations professional with 10 years of experience supporting enterprise-scale AWS environments, production operations, and mission-critical workloads. Currently leading a 15-member AWS Infrastructure Support team managing 800+ AWS accounts, responsible for incident management, platform reliability, production support, stakeholder collaboration, and service delivery. Hands-on experience with AWS, Linux, Terraform, Infrastructure as Code (IaC), DevOps practices, cloud automation, Big Data, and the Hadoop ecosystem. Proven ability to resolve critical production incidents, drive operational improvements, and maintain high service availability and SLA compliance. AWS Certified Solutions Architect – Associate, AWS Certified Developer – Associate, and AWS Certified Data Engineer – Associate.

What I'm looking for

I'm seeking opportunities as a **Cloud Infrastructure Team Lead, AWS Cloud Engineer, Site Reliability Engineer (SRE), Platform Engineer, or DevOps Engineer**, where I can leverage my 10 years of experience in AWS cloud operations, production support, infrastructure automation, and team leadership. I'm passionate about building reliable, secure, and scalable cloud platforms while driving operational excellence through automation, Infrastructure as Code (Terraform), and DevOps best practices.

Experience

Accenture

Team Lead, Cloud Infrastructure Operations

Accenture

Dec 2021 – Present

**Project: SRE Cloud Infrastructure Support | Client: Travelers Insurance** **Jul 2022 – Present** * Lead the AWS Infrastructure Support team managing **800+ AWS accounts**, ensuring platform reliability, incident management, production support, and operational excellence. * Manage AWS infrastructure operations, including provisioning, deprovisioning, IAM access management, governance, and Terraform automation. * Troubleshoot AWS services including EC2, IAM, VPC, Route 53, EBS, CloudWatch, Auto Scaling, and ELB to maintain high availability. * Lead P1/P2 incident resolution, root cause analysis (RCA), and problem management while ensuring SLA compliance. * Monitor infrastructure using Amazon CloudWatch and perform Linux administration, troubleshooting, patch management, and performance analysis. * Coordinate production releases, collaborate with cross-functional teams, mentor engineers, and standardize operational processes. * Provision AWS infrastructure using Terraform, including EC2, IAM, Security Groups, VPCs, and AWS account configurations. **Key Achievements** * Promoted to Team Lead, leading cloud operations across **800+ AWS accounts**. * Resolved **100+ P1/P2 incidents** while meeting SLA commitments. * Improved platform reliability and streamlined support processes. * Conducted **20+ knowledge transfer sessions** to accelerate onboarding. **Project: BlackRock Aladdin Wealth Data Platform | Client: JPMorgan Chase & Co.** **Dec 2021 – Jul 2022** * Developed AWS data pipelines using Python, PySpark, Pandas, and SQL. * Built ETL workflows using AWS Glue and managed Hive, Athena, and Amazon Redshift. * Automated workflows using Control-M and performed capacity planning, performance tuning, and resource optimization. * Provided production support for AWS data platforms, collaborating with development and DevOps teams to ensure platform reliability.

aws cloudTerraformlinuxcontrol-mService-Now
Yash Technologies

Software Engineer

Yash Technologies

Sep 2016 – Dec 2021

**Azure HDInsight Hadoop Platform | Client: Caterpillar Inc., USA** **Aug 2020 – Dec 2021** * Administered Azure HDInsight Hadoop clusters, managing platform lifecycle activities including provisioning, upgrades, service operations, and compatibility validation. * Resolved infrastructure and performance issues in collaboration with Microsoft Support and DevOps teams, ensuring platform availability and operational stability. * Applied security patches and compliance updates to maintain platform security and meet organizational standards. * Designed and implemented a shared cluster strategy, improving resource utilization and reducing infrastructure costs. **Environment:** Azure HDInsight, Hadoop, Hive, YARN, HDFS, Linux --- **Hadoop Development & IoT Data Ingestion | Client: Caterpillar Inc., USA** **Sep 2016 – Aug 2020** * Configured, administered, and monitored multi-tenant Hadoop clusters supporting enterprise applications and large-scale data workloads. * Performed Hadoop CDH patching, version upgrades, and platform maintenance to ensure stability, security, and high availability. * Administered Hadoop ecosystem components including Hive, Sqoop, Flume, Oozie, and YARN, and automated recurring operational tasks using Oozie workflows. * Developed and supported IoT data ingestion pipelines, integrating sensor data (SC2-SuperComm2, VIMS) into Hadoop and downstream platforms including Teradata, HBase, and OpenTSDB. * Monitored platform health, managed backups, ensured data integrity, and provided operational metrics and status reports to stakeholders. **Environment:** Hadoop, Hive, Sqoop, Flume, Oozie, YARN, HBase, Teradata, OpenTSDB, Linux

Hadoop AdministrationHiveclouderaHbaseFlume

Education

Centre for Development of Advanced Computing (CDAC), Pune

PG Diploma

2016 – 2016

Lakshmi Narain College of Technology and Science, Bhopal

B.Tech / B.E. · CSE

2011 – 2015

Certifications

GitHub Copilot

Microsoft

Oct 2025 – Oct 2027

Skills

SQLPythonGithubTerraformHadoop AdministrationCI/CDLinuxAWS Cloud

Languages

Hindi (Professional working proficiency)English (Professional working proficiency)Marathi (Native or bilingual proficiency)