Job description
- Lead infrastructure architecture, capacity planning, scalability, redundancy, and disaster recovery.
- Lead production incident response, root cause analysis, and preventive actions.
- Own infrastructure and network security, including WAF, DDoS mitigation, and security findings.
- Drive cloud cost optimization through resource monitoring and efficient infrastructure utilization.
- Prepare and present architectural decisions and RCAs for technical and executive audiences.
- Manage infrastructure and security vendors, including contract scope and SLA performance.
- Maintain infrastructure documentation, including runbooks, architecture diagrams, and configuration standards.
- Mentor and provide technical guidance to the Infrastructure Team.
Job requirements
- Strong hands-on experience with AWS infrastructure, including EC2, RDS, ECS Fargate, ElastiCache, CloudFront, WAF, ALB, and VPC.
- Strong understanding of cloud architecture, scalability, capacity planning, security, and disaster recovery.
- Experience with Infrastructure as Code (IaC) and CI/CD, such as Terraform or CloudFormation.
- Strong troubleshooting skills with experience leading production incidents and root cause analysis.
- Data-driven approach to cloud cost and infrastructure performance optimization.
- Strong problem-solving and architectural decision-making skills.
- Ability to mentor and provide technical guidance to the Infrastructure Team.