| Job Title | Manager – ITSM – Engineering – Platform Development & Integration |
| Company | Deloitte Touche Tohmatsu India LLP |
| Job Requisition ID | 109995 |
| Location | Bengaluru, Karnataka |
| Job Type | Full Time |
| Department/Domain | Engineering – Platform Development & Integration / Cloud Operations |
| Role | Lead Production Incident Manager – Enterprise Technology & Performance |
| Experience | 10–15+ years overall IT experience; 8+ years in Enterprise Production Support, Incident Management, Cloud Infrastructure & SRE |
| Qualification | Bachelor’s or Master’s degree in Computer Science, IT, Engineering or related field |
| Domain Experience | Banking & Financial Services preferred, especially Cards & Payments, Mobile Applications and Cloud-Native solutions |
| Incident Management | Lead P1–P4 incidents, major incident management, escalation, communication, resolution and closure |
| ITIL | Incident, Problem & Change Management following ITIL best practices |
| Production Support | Lead 24×7 L2/L3 production support for enterprise applications and cloud infrastructure |
| Cloud | AWS – EC2, S3, RDS, Lambda, VPC, IAM, DynamoDB, CloudWatch |
| SRE | Automation, monitoring, observability, reliability, MTTR reduction and operational excellence |
| Infrastructure | Linux, networking, load balancing, cloud security, SQL/NoSQL databases and REST APIs |
| Containers | Docker and Kubernetes for deployment, scaling and workload orchestration |
| DevOps | Jenkins, CI/CD, Terraform, CloudFormation, Ansible, Python and Bash |
| Monitoring Tools | ELK, Kibana, Grafana, CloudWatch, Splunk, Prometheus, Nagios, Zenduty and Site24x7 |
| Disaster Recovery | DR planning, Active-Active/Active-Passive failover and High Availability architecture |
| RCA & Service Recovery | Root Cause Analysis, Post Incident Reviews, CAPA, service recovery and corrective actions |
| Database & API Support | Complex SQL DDL/DML queries, REST API validation, cache management and health checks |
| Team Leadership | Manage and mentor SRE and Production Support teams; ensure SLA, KPI, SLO and MTTR compliance |
| Stakeholder Management | Coordinate with technical teams, business stakeholders and clients during critical incidents |
| Optimization | Cloud infrastructure optimization, automation, capacity planning and cost optimization |
| Customer Experience | Monitor CSAT/NPS and drive continuous service improvement |
| Preferred Certifications | ITIL Foundation; AWS Certified Solutions Architect – Associate/Professional highly preferred |
| Key Skills | ITSM, AWS, SRE, Incident Management, Cloud Infrastructure, Kubernetes, Docker, DevOps, ITIL, Production Support, Disaster Recovery, Automation, Monitoring |