Site Reliability Engineer

Finthrive · Gurgaon

  • Experience6–10 yrs
  • SalaryNot disclosed
  • Work modehybrid
  • Levelsenior
  • Posted15 Sept 2026

About Finthrive

Finthrive is hiring in Gurgaon in healthcare. This role looks for around 6+ years of experience.

Skills

  • Microsoft Azure
  • Azure App Services
  • Azure Application Gateway
  • Azure Front Door
  • Terraform
  • Azure Bicep
  • ARM Templates
  • Azure Automation
  • Azure Functions
  • Azure Monitor
  • Application Insights
  • Log Analytics
  • Kusto Query Language
  • Grafana
  • Incident Management
  • Root Cause Analysis
  • Change Management
  • SLA
  • SLO
  • Error Budget
  • Capacity Planning
  • CI/CD
  • Azure DevOps
  • REST
  • Postman
  • SoapUI
  • Version Control
  • Release Management
  • Bachelor's Degree

The role

A site reliability engineer at a healthcare technology company designs and operates Azure cloud platforms, strengthening distributed-system resilience through infrastructure as code, observability, and incident management. The role also applies Kusto Query Language and Terraform to automate operations and optimize application performance.

Full job description

Role & responsibilities

Site Reliability Engineer / Cloud Engineer

SRE & Reliability Engineering

Managed production environments ensuring high availability and reliability of cloud-hosted applications

Led incident response, performed deep root cause analysis, and implemented preventive measures to reduce recurrence

Improved system resilience through proactive monitoring and performance tuning strategies

Azure Application & Platform Engineering

Designed and supported application architectures using:

Azure App Services and App Service Plans

Azure App Service Environment v3 (ASEv3) for isolated, high-scale workloads

Azure Application Gateway (WAF-enabled) for L7 traffic management

Azure Front Door for global traffic routing and failover

Implemented secure and scalable cloud networking patterns, optimizing latency and throughput

Automation & Toil Reduction

Identified repetitive operational tasks and reduced manual effort through automation-first solutions

Developed automation using:

Terraform / Bicep / ARM templates

Azure Automation (Hybrid Workers)

Azure Functions for event-driven workflows

Leveraged AI-assisted tools (GitHub Copilot, Copilot) to accelerate scripting and automation development, while ensuring strict validation for enterprise use

Observability & Monitoring

Built and enhanced observability using:

Azure Monitor, Application Insights, Log Analytics

Created KQL-based queries and dashboards for proactive issue detection

Reduced false alerts by optimizing alert thresholds and improving signal quality

Performance & System Optimization

Analyzed application performance across distributed systems to identify bottlenecks

Implemented improvements through:

Scaling strategies (horizontal & vertical)

Network optimization (AGW / Front Door tuning)

Backend service improvements

Collaboration & Engineering Enablement

Partnered with SRE, CloudOps, and development teams to design resilient systems

Contributed to runbooks, documentation, and operational standards

Enabled engineering teams by improving platform reliability and deployment pipelines

Key Achievements

Reduced manual operational effort by X% through automation initiatives

Improved system availability to 99.X% by strengthening monitoring and failure handling mechanisms

Decreased incident resolution time by X% via enhanced observability and streamlined runbooks

Optimized application performance using Front Door and AGW tuning, reducing latency by X%

Preferred candidate profile

Cloud & Platform Engineering

Microsoft Azure (Preferred)

Understanding and experience in developing Azure function Apps, Azure logic Apps

Understanding of event triggers, event hub, service bus.

Azure Landing zones, Azure Cloud Adoption Framework, Azure Well Architectured Framework

Application Hosting: App Services, App Service Plans, ASEv3

Networking: Azure Application Gateway (AGW), Azure Front Door, VNet, NSGs, Load Balancing

Cloud Architecture: High Availability, Fault Tolerance, Scalability Patterns

Incident Management and RCA

Incident Management, P1 troubleshooting, Change Management

Experienced in leading RCA and representing on the weekly call

SLA / SLO / Error Budget concepts

System Performance Optimization & Capacity Planning

Toil Reduction through Automation

Infrastructure as Code & Automation

Terraform, Azure Bicep, ARM Templates

Azure Automation (Hybrid Workers)

Azure Functions (Serverless automation)

API-based automation and orchestration

Observability & Monitoring

Azure Monitor, Log Analytics Workspace, Grafana, Site 24x7 (or similar SaaS based synthetic monitoring tool)

Application Insights

KQL (Kusto Query Language)

Alert tuning and signal-to-noise optimization

AI-Enabled Productivity (Not as Skill)

Leveraging GitHub Copilot / Microsoft Copilot for:

Code acceleration and script generation

Automation development support

Troubleshooting and log analysis assistance

Proven track record of workforce optimization leveraging AI tools.

Applying validation frameworks to ensure secure, accurate, and production-grade outputs

DevOps & Integration

CI/CD using Azure DevOps

Deep understanding on version control

API integrations (REST, Postman, SoapUI)

Source control and release management

Bachelors Degree in Computer Science / Engineering or related field

Preferred/Additional Experience

Experience with microservices and distributed architectures

Exposure to low-code automation platforms

Working knowledge of AWS cloud services

Preferred/Additional Certifications

AZ-104 Azure Administrator

AZ-700 Designing and Implementing Microsoft Azure Networking Solutions

AZ-400 Microsoft Certified: DevOps Engineer Expert