CloudOps Engineer IV

PDI Technologies · Hyderabad

  • Experience10–11 yrs
  • SalaryNot disclosed
  • Work modehybrid
  • Posted28 Sept 2026

About PDI Technologies

PDI Technologies is hiring in Hyderabad in ecommerce retail. This role looks for around 10+ years of experience.

Skills

  • AWS
  • Azure
  • Kubernetes
  • Helm
  • Argo CD
  • Argo Workflows
  • Terraform
  • OpenTofu
  • Jenkins
  • Rancher
  • Datadog
  • Site Reliability Engineering

The role

A site reliability engineer at a payments and retail technology company builds and operates AWS and Azure infrastructure, Kubernetes clusters, and GitOps delivery. This person strengthens Terraform, Argo CD, and Datadog practices for reliable customer-facing platforms.

Full job description

PDI Technologies is looking for a CloudOps Engineer IV to join the SRE organization supporting Paylo, PDI’s payments, loyalty, and fuel-pricing product suite. This is a senior, hands-on individual-contributor role focused on keeping high-traffic, customer- and partner-facing platforms reliable, secure, and running efficiently across multi-cloud infrastructure.

You will bring strong, hands-on experience across AWS, Azure, Kubernetes, Helm, Argo CD, Terraform/OpenTofu, Jenkins, and Datadog, and apply it directly — building and operating infrastructure, improving deployment pipelines, and strengthening observability. You will work under the direction of the Senior Manager, SRE, executing against the team’s reliability and infrastructure roadmap while bringing your own judgment and technical leadership to the problems in front of you.

Job Responsibilities: Cloud Infrastructure & Operations:

Build, operate, and troubleshoot infrastructure across AWS and Azure in support of production workloads.

Operate and maintain Kubernetes clusters, including deploying and maintaining Helm charts for the services you support.

Participate in on-call rotation, respond to incidents, and drive them to resolution within your area of ownership

Contribute to capacity planning, cost optimization, and resilience improvements for the systems you support

Automation & Continuous Delivery:

Build and maintain GitOps-based deployment pipelines using Argo CD/Argo Workflows, including rollout and promotion configuration across environments.

Write and maintain Infrastructure-as-Code (Terraform, OpenTofu) for the infrastructure you own, following team module standards.

Build and maintain CI/CD pipelines in Jenkins, improving build/deploy automation and reliability.

Support progressive delivery practices (blue-green/canary, automated rollback) for the services you support

Reliability & Observability:

Build and maintain Datadog dashboards, monitors, and alerts for the services you support, tuning alert thresholds to reduce noise.

Contribute to defining SLIs/SLOs for your services and help track them over time.

Participate in postmortems for incidents you're involved in, and follow through on assigned remediation items

Collaboration & Mentorship:

Partner with engineers across the SRE team and with product engineering teams to troubleshoot issues and improve system design.

Share knowledge with and mentor less-experienced engineers on the team (CloudOps Engineer I III) on cloud infrastructure, Kubernetes, and CI/CD practices.

Contribute to documentation, runbooks, and onboarding materials for the systems you support.

Required Qualifications:

10+ years of experience in Cloud Operations, Site Reliability Engineering, DevOps, or Infrastructure Engineering roles

Hands-on experience with AWS — you can build, troubleshoot, and operate cloud infrastructure directly.

Hands-on experience with Kubernetes and Helm — deploying, operating, and troubleshooting workloads in production clusters.

Hands-on experience with Argo CD/Argo Workflows for GitOps-based continuous delivery.

Hands-on experience with Infrastructure as Code (Terraform, OpenTofu).

Hands-on experience with Jenkins and Rancher for CI/CD pipeline development and maintenance.

Hands-on experience with Datadog (or equivalent observability platform), including building dashboards, monitors, and alerts.

Experience participating in an on-call rotation and responding to production incidents.

Strong communication skills and the ability to work effectively across teams.

Preferred Qualifications:

Experience supporting payments, fuel/retail, or loyalty platforms, or other systems with PCI DSS or similar compliance obligations

Relevant certifications such as CKA/CKAD, AWS Certified Solutions Architect – Associate, Microsoft Certified: Azure Administrator, or HashiCorp Terraform Associate.

Experience with messaging systems (Kafka/SQS/SNS) and multi-region/multi-AZ resilience patterns.

Prior experience mentoring junior engineers or leading small technical initiatives.

Behavioral Competencies:

Cultivates Innovation

Decision Quality

Manages Complexity

Drives Results

Business Insight

PDI is committed to offering a well-rounded benefits program, designed to support and care for you, and your family throughout your life and career.  This includes a competitive salary, market-competitive benefits, and a quarterly perks program. We encourage a good work-life balance with ample time off [time away] and, where appropriate, hybrid working arrangements.  Employees have access to continuous learning, professional certifications, and leadership development opportunities. Our global culture fosters diversity, inclusion, and values authenticity, trust, curiosity, and diversity of thought, ensuring a supportive environment for all.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.