Role: DevOps Engineer III
Experience: 6–9 Years
Employment Type: Full-Time
Function: Engineering / Infrastructure & Platform
About POP
We’re POP, a new-age fintech company that’s making payments fun, rewarding, and totally worth it! With POP UPI, you can pay anyone, anywhere and earn POPcoins every time you do. Those POPcoins aren’t just points, they're real rewards you can use to shop, save, and score deals from over 500+ awesome brands. Founded in Bengaluru in 2023, we’re backed by Razorpay and partnered with YES Bank to bring India’s most rewarding payment experience. At POP, we believe every payment should give something back so we’re on a mission to make spending smarter, shopping more exciting, and rewards a part of everyone’s daily life.
About the Role
We are looking for a seasoned DevOps Engineer III to join our growing engineering team. In this role, you will be a key contributor and technical anchor for our cloud infrastructure, automation, and platform reliability initiatives. You will collaborate closely with development, security, and product teams to ensure scalable, secure, and highly available systems — while also mentoring and managing the day-to-day priorities of the DevOps team.
Key Responsibilities
Cloud Infrastructure & Architecture
- Design, provision, and manage cloud infrastructure on AWS, with a strong focus on services including EKS, IAM, S3, RDS, Lambda, VPC, CloudWatch, ECR, Route53, SQS, SNS, MSK, MongoDB, ElasticCache, API Gateway(Kong) and related services.
- Manage and optimize Amazon EKS clusters including node groups, autoscaling, namespaces, RBAC, and workload deployments.
- Architect and enforce IAM policies, roles, and permission boundaries following the principle of least privilege.
- Manage RDS instances (MySQL/PostgreSQL) including backups, parameter groups, replication, and failover configurations.
- Design and maintain S3 bucket policies, lifecycle rules, versioning, and cross-account access patterns.
- Build and maintain Lambda-based serverless workflows and event-driven architectures.
Cost Optimization & FinOps
- Own and drive cloud cost optimization across all infrastructure — continuously monitor, analyze, and reduce AWS spend without compromising reliability or performance.
- Implement FinOps practices including tagging strategies, cost allocation, chargeback/showback models, and budget alerting using AWS Cost Explorer, Budgets, and Trusted Advisor.
- Identify and act on cost-saving opportunities such as Reserved Instances, Savings Plans, Spot Instances, and right-sizing of compute, storage, and database resources.
- Optimize EKS workloads through efficient bin-packing, autoscaling (Cluster Autoscaler / KEDA), and node group rightsizing.
- Audit and eliminate idle or underutilized resources — unused EBS volumes, unattached EIPs, over-provisioned RDS instances, and stale snapshots.
- Enforce cost-aware engineering practices within the team — reviewing infrastructure PRs for cost impact and promoting efficient architecture patterns.
- Publish regular cost dashboards and savings reports to engineering and leadership stakeholders.
- Evaluate and adopt cost-efficient architectural patterns such as serverless, event-driven, and managed services where applicable.
Infrastructure Automation & IaC
- Own and drive end-to-end infrastructure lifecycle management using Terraform — from module development, state management, and workspace strategy to plan/apply pipelines.
- Maintain Terraform codebase with best practices: modularization, remote backends (S3 + DynamoDB), locking, and environment-based variable management.
- Automate provisioning, scaling, patching, and decommissioning of infrastructure components.
- Build and maintain CI/CD pipelines using tools like GitHub Actions, Jenkins, or GitLab CI for application and infrastructure deployments.
Logging, Monitoring & Incident Management
- Implement and manage observability stacks covering metrics, logs, and distributed tracing.
- Work with Datadog for infrastructure and application monitoring — dashboards, APM, log management, synthetic monitoring, and alerting.
- Configure and manage Zenduty for on-call schedules, escalation policies, and incident routing.
- Define and track SLOs/SLAs, set up anomaly detection, and drive incident postmortems.
- Manage centralized logging pipelines using tools such as ELK/OpenSearch, Fluent Bit, or CloudWatch Logs Insights.
Team Leadership & Collaboration
- Act as a technical lead for the DevOps team — facilitating sprint planning, task prioritization, and workload balancing.
- Mentor junior and mid-level DevOps engineers through code reviews, pair programming, and knowledge-sharing sessions.
- Collaborate with development teams on architecture decisions, deployment strategies, and capacity planning.
- Maintain and improve runbooks, SOPs, and infrastructure documentation.
- Drive adoption of DevOps best practices across engineering teams.
Business Continuity & Disaster Recovery
- Hands-on experience in planning, executing, and documenting Business Continuity Plan (BCP) and Disaster Recovery (DR) drills.
- Define and maintain RTO (Recovery Time Objective) and RPO (Recovery Point Objective) targets for critical systems and services.
- Design and implement DR strategies including multi-region failover, database replication, automated backups, and infrastructure redundancy on AWS.
- Conduct periodic DR drills and tabletop exercises — simulating failure scenarios such as region outages, data corruption, or service degradation — and document outcomes and action items.
- Maintain and continuously improve runbooks and playbooks for disaster recovery and business continuity scenarios.
- Collaborate with business and product stakeholders to identify critical systems and define recovery priorities.
- Ensure DR readiness is validated through automated testing and regular review cycles.
Security, Compliance & Audits
- Play an active role in security audits and compliance programs, including but not limited to:
- ISO 27001
- PCI-DSS
- SOC 2 (Type I & II)
- External Audits & Assessments
- ASV (Approved Scanning Vendor) Scans
- VAPT (Vulnerability Assessment & Penetration Testing)
- Implement and track remediation of findings from security assessments and audit reports.
- Ensure infrastructure compliance through policy-as-code, automated scanning (e.g., Checkov, tfsec, AWS Config, GuardDuty, SecurityHub).
- Manage secrets and sensitive configurations using AWS Secrets Manager or
HashiCorp Vault.
- Participate in audit evidence collection, documentation, and liaison with auditors.
AI Adoption & AI-First Engineering
- Champion an AI-first mindset across DevOps and infrastructure engineering — proactively identifying opportunities to integrate AI/ML tools into workflows, pipelines, and operations.
- Leverage AI-powered tools to enhance developer productivity, incident response, and infrastructure management — including AI coding assistants, AIOps platforms, and intelligent alerting systems.
- Explore and implement AWS AI/ML services (e.g., Amazon Bedrock, SageMaker, CodeWhisperer) to automate repetitive operational tasks and improve system intelligence.
- Drive adoption of AI-driven observability and anomaly detection within Datadog or equivalent platforms to proactively surface issues before they impact users.
- Collaborate with product and data engineering teams to support the infrastructure needs of AI/ML workloads — including GPU instances, high-throughput storage, and model serving infrastructure.
- Stay current with the rapidly evolving AI tooling landscape and bring relevant innovations back to the team through demos, POCs, and knowledge-sharing sessions.
- Promote the use of AI-assisted automation in areas such as auto-remediation, predictive scaling, intelligent log analysis, and security threat detection.
Required Skills & Qualifications
- 6–9 years of overall experience in DevOps / Cloud / Infrastructure engineering.
- Strong hands-on expertise in AWS — EKS, IAM, S3, RDS, Lambda, VPC, CloudWatch, and other core services.
- Deep proficiency in Terraform — modules, workspaces, remote state, and E2E infrastructure lifecycle management.
- Solid experience with Kubernetes — cluster administration, Helm charts, ingress controllers, and workload management.
- Proven experience with CI/CD pipelines (GitHub Actions / Jenkins / GitLab CI or equivalent).
- Hands-on experience with Datadog and Zenduty (or equivalent monitoring and incident management platforms).
- Experience participating in or leading security compliance programs (ISO, PCI-DSS, SOC2, VAPT, ASV).
- Demonstrated experience in BCP/DR planning and drill execution, with the ability to present findings and remediation plans to stakeholders.
- Strong scripting skills in Python, Bash, or equivalent.
- Comfortable working with Linux environments, networking fundamentals, and container technologies (Docker).
- Excellent communication skills and the ability to work cross-functionally.
Good to Have
- Experience with GitOps workflows (ArgoCD / Flux).
- Exposure to cost optimization strategies in AWS (Reserved Instances, Savings Plans, right-sizing).
- Familiarity with service mesh tools like Istio or Linkerd.
- Knowledge of chaos engineering practices.
- Experience with multi-account AWS Organizations and landing zone setups.
- AWS certifications (Solutions Architect, DevOps Professional, Security Specialty).