Site Reliability Engineer (SRE I)

Leena AI · Gurgaon

  • Experience1–3 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Leveljunior
  • Posted11 Sept 2026

About Leena AI

Leena AI is hiring in Gurgaon in technology software. This role looks for around 1+ years of experience.

Skills

  • Linux
  • Networking
  • Kubernetes
  • Docker
  • Amazon Web Services
  • Microsoft Azure
  • Google Cloud Platform
  • Helm
  • Git
  • Prometheus
  • Grafana
  • Tempo
  • Bash
  • Python

The role

A site reliability engineer at an enterprise AI software company monitors and automates production systems using Kubernetes, Docker, and observability practices, while improving cloud infrastructure and incident response. The role also applies Terraform and Prometheus to strengthen scalable platform operations.

Full job description

About Leena AI

Leena AI is a leader in Agentic AI for the enterprise. We are building an iconic company, delivering AI Colleagues that transform back-office functions and accelerate the full promise of Generative AI—unlocking real productivity gains, cutting costs, and delighting employees at scale.

Leena AI provides the most forward-looking, open, and scalable Agentic AI architecture for the enterprise— it empowers CIOs and CTOs to develop, deploy, and manage AI Colleagues for the back office at scale. Built with full governance, compliance, security, and auditability at its core.

Leena AI integrates with 1000+ applications, including SAP, Salesforce, ServiceNow, Workday, and Microsoft Office 365. We are proud to be trusted by 500+ global enterprises and 20 million+ employees, including leading brands such as Nestlé, Puma, Coca-Cola, Sony, and Etihad Airways.

Founded in 2018 and headquartered in New York, Leena AI has secured over $40M in financing from top-tier investors including Greycroft, Bessemer Venture Partners, B Capital, and Y Combinator.

About The Role

We're hiring an SRE I to help keep our production systems reliable, observable, and scalable. You'll workalongside the team to monitor systems, automate toil, and respond to incidents.

What You'll Do

Monitor production systems and respond to alerts; join the on-call rotation. Help track reliability metrics (SLIs/SLOs) and reduce toil through automation. Assist with deployments, rollbacks, and CI/CD maintenance. Investigate incidents and contribute to postmortems. Improve dashboards, alerting, and documentation.

Skills & Tools

Working knowledge of Linux and networking fundamentals. Familiarity with a cloud provider (AWS, Azure, or GCP). Hands-on Kubernetes and Docker — comfortable debugging pods (e.g. OOMKills, scheduling failures),

not just deploying apps.

Exposure to Terraform (or similar IaC), Helm, and Git. Monitoring and troubleshooting with Prometheus, Grafana, and Tempo (or similar). Scripting in any language (e.g. Bash, Python).

What We're Looking For

1-3 years in SRE/DevOps/systems, or a strong internship/project background. Curiosity, eagerness to learn, and a methodical approach to troubleshooting.

Skills: kubernetes,enterprise,aws,reliability,docker,azure