Analyst II, Production Support

FIS · Pune/Pimpri-Chinchwad Area

  • Experience3–4 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Levelmid
  • Posted23 Sept 2026

About FIS

FIS is hiring in Pune/Pimpri-Chinchwad Area in financial services. This role looks for around 3+ years of experience.

Skills

  • AWS
  • Microsoft Azure
  • Linux/UNIX
  • Kubernetes
  • Docker
  • Terraform
  • CI/CD
  • Python
  • PowerShell
  • Bash
  • Observability
  • Incident Management
  • SQL Server
  • Cloud Security
  • ServiceNow

The role

A site reliability and cloud engineer in a financial services company designs and operates production services on AWS or Microsoft Azure. They automate infrastructure with Terraform and Python, run Kubernetes workloads, and improve reliability through observability, incident response, and resilient architecture. Their skills include AWS, Microsoft Azure, Terraform, Kubernetes, Python, Linux, CI/CD, ServiceNow, and cloud security.

Full job description

Position Type

Full time

Type Of Hire

Experienced (relevant combo of work and education)

SITE RELIABILITY & CLOUD ENGINEER

Production Engineering | Pune, India

ROLE PURPOSE

Engineer reliable, secure and scalable cloud services by combining production operations with software engineering, automation, observability and cloud platform expertise. The role owns service reliability outcomes across the application lifecycle and partners with development, infrastructure, security and support teams to prevent incidents, reduce toil and improve customer experience.

Role Overview

We are seeking an experienced Site Reliability and Cloud Engineer with strong hands-on expertise in AWS and/or Microsoft Azure. The successful candidate will design, deploy, operate and continuously improve production services, with a focus on availability, performance, resilience, security and operational efficiency.

This is an engineering-led production role. The engineer will use automation, infrastructure as code, telemetry and disciplined incident/problem management to create sustainable services, accelerate safe change and reduce manual operational effort.

Key Responsibilities

Own the reliability, availability, performance and operational readiness of cloud-hosted applications and platform servicesDefine and maintain service level indicators (SLIs), service level objectives (SLOs), availability targets and actionable service-health dashboardsBuild monitoring and alerting around customer-impacting symptoms, golden signals and service dependencies; reduce alert noise and improve diagnostic qualityAutomate repeatable operational work, remediation, deployments, configuration, evidence collection and environment validation using code and pipelinesDesign, build and maintain cloud infrastructure using infrastructure as code, reusable modules, policy guardrails and secure engineering standardsParticipate in incident response and on-call support, including rapid triage, stabilisation, technical escalation and clear stakeholder communicationLead or contribute to blameless post-incident reviews; identify root causes, track corrective actions and engineer controls that prevent recurrencePartner with application engineering teams on architecture, capacity planning, performance engineering, release readiness and production operabilityImprove deployment safety through CI/CD controls, automated testing, progressive validation, rollback strategies and change-risk reductionEngineer resilience through high-availability patterns, backup and restore validation, disaster-recovery runbooks, failover testing and dependency mappingManage production risks including vulnerabilities, patching, certificates, secrets, access controls, audit evidence and cloud security findingsCreate and maintain runbooks, standard operating procedures, architecture records and operational knowledge that support consistent 24x7 service deliveryAnalyse operational data and trends to reduce mean time to detect and restore service, eliminate recurring failure modes and improve capacity and cost efficiencyMentor support and engineering colleagues in SRE practices, automation, observability, troubleshooting and operational ownership

Required Skills And Experience

Three or more years of experience in cloud engineering, production engineering, DevOps, platform engineering or Site Reliability EngineeringHands-on experience operating production workloads on AWS and/or Microsoft Azure, including compute, networking, identity, storage, managed databases and monitoring servicesStrong Linux/UNIX administration and troubleshooting skills; working knowledge of Windows Server is beneficialPractical experience with Kubernetes and containers, including EKS and/or AKS, Docker, deployment troubleshooting and workload reliabilityInfrastructure-as-code experience using Terraform; ability to build reusable, governed and maintainable modulesCI/CD experience with tools such as Harness, Azure DevOps, GitHub Actions, Jenkins or equivalent, including deployment and rollback controlsProgramming or advanced automation capability using Python, PowerShell, Bash or a comparable language; coding experience beyond simple one-off scriptsExperience with observability platforms and practices covering metrics, logs, traces, dashboards, alerting and application performance monitoringStrong incident and problem management experience, including technical triage, root cause analysis, corrective actions and production communicationsUnderstanding of distributed systems, scalability, high availability, capacity management, performance bottlenecks and failure modesExperience supporting SQL Server and cloud database services, including connectivity, performance diagnostics, backup/restore and operational monitoringWorking knowledge of cloud security, least privilege, certificate and secrets management, vulnerability remediation, auditing and compliance controlsExperience with ServiceNow or a comparable IT service management platform for incidents, problems, changes and operational work trackingClear written and verbal communication, disciplined documentation, strong ownership and the ability to work across engineering and business teams

Desirable Skills

Cloud certification in AWS or Microsoft Azure; Kubernetes or Terraform certification is advantageousExperience with Akamai, API gateways, web application delivery, DNS, load balancing, VPNs and enterprise network connectivityKnowledge of event-driven and service-oriented architectures, domain-driven design and messaging platformsExperience in regulated financial services or another environment with formal change, risk, audit, resilience and data-protection obligationsExperience implementing policy as code, security scanning, automated compliance controls, FinOps or cloud cost optimisationExperience supporting globally distributed services, customer onboarding and follow-the-sun operational models

Success Measures

Outcome

Evidence of Success

Reliability

Services have defined health measures, meaningful alerts, tested recovery procedures and improving availability trends.

Engineering efficiency

Manual toil and recurring incidents are reduced through automation, reusable tooling and preventive engineering.

Operational readiness

Releases and services meet documented production-readiness, security, observability and support requirements.

Incident learning

Major incidents produce clear root causes, owned corrective actions and measurable risk reduction.

Collaboration

Development, platform, security and support teams share clear ownership and use consistent operational practices.

Role Expectations

Location~ Pune, IndiaFlexible to work on rotational shifts and weekends tooWork closely with global engineering, production support, security and service-management teamsParticipate in an agreed on-call or out-of-hours support rotation where required for production servicesDemonstrate an engineering mindset~ automate where practical, design for failure, measure service health and treat operational learning as product improvement

Privacy Statement

FIS is committed to protecting the privacy and security of all personal information that we process in order to provide services to our clients. For specific information on how FIS protects personal information online, please see the Online Privacy Notice.

Sourcing Model

Recruitment at FIS works primarily on a direct sourcing model; a relatively small portion of our hiring is through recruitment agencies. FIS does not accept resumes from recruitment agencies which are not on the preferred supplier list and is not responsible for any related fees for resumes submitted to job postings, our employees, or any other part of our company.

#pridepass