Engineering Manager - DevOps

Leena AI · Gurgaon

  • Experience8–12 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Levelexecutive
  • Posted11 Sept 2026

About Leena AI

Leena AI is hiring in Gurgaon in technology software. This role looks for around 8+ years of experience.

Skills

  • Kubernetes
  • Terraform
  • Helm
  • Amazon Web Services (AWS)
  • Microsoft Azure
  • Google Cloud Platform (GCP)
  • Python
  • Go
  • Bash
  • CI/CD
  • GitHub Actions
  • GitLab CI
  • ArgoCD
  • Prometheus
  • Grafana
  • ELK
  • OpenTelemetry
  • HashiCorp Vault
  • Key Management Service (KMS)
  • Software Bill of Materials (SBOM)
  • CIS Benchmarks
  • SOC 2
  • ISO 27001
  • Infrastructure as Code
  • Cloud cost management
  • Secrets management
  • Image signing
  • Vulnerability management
  • Incident response
  • Postmortems
  • MongoDB
  • Elasticsearch
  • Redis
  • RabbitMQ
  • ClickHouse

The role

An engineering manager at an agentic AI software company owns multi-cloud infrastructure and Kubernetes, building secure private-cloud deployments with Terraform and Helm. Cost optimization, observability operations, and infrastructure security shape reliable enterprise platforms.

Full job description

About Leena AI

Leena AI is a leader in Agentic AI for the enterprise. We are building an iconic company, delivering AI Colleagues that transform back-office functions and accelerate the full promise of Generative AI - unlocking real productivity gains, cutting costs, and delighting employees at scale.

Leena AI provides the most forward-looking, open, and scalable Agentic AI architecture for the enterprise - it empowers CIOs and CTOs to develop, deploy, and manage AI Colleagues for the back office at scale. Built with full governance, compliance, security, and auditability at its core.

Leena AI integrates with 1000+ applications, including SAP, Salesforce, ServiceNow, Workday, and Microsoft Office 365. We are proud to be trusted by 500+ global enterprises and 20 million+ employees, including leading brands such as Nestlé, Puma, Coca-Cola, Sony, and Etihad Airways.

Founded in 2018 and headquartered in New York, Leena AI has secured over $40M in financing from top-tier investors including Greycroft, Bessemer Venture Partners, B Capital, and Y Combinator.

Role Overview

Leena AI runs a multi-tenant agentic AI platform for large global enterprises across fifteen production regions on AWS, Azure and GCP, plus private-cloud deployments inside customer environments. This role owns the infrastructure all of that runs on: Kubernetes clusters, cloud accounts, networking, CI runners, the observability stack, secrets and supply-chain security, and the cloud bill.

You lead a team of 3–5 DevOps engineers & SREs. You work alongside our Platform Engineering group, which owns service-level objectives, alerting and incident command for the product; your team is the infrastructure counterpart that keeps the ground solid and makes it cheaper and safer every quarter.

This is a hands-on management role. You will still write Terraform and review Helm charts. You will also carry a second product: our deployment package for private-cloud and OEM partners (Terraform, Helm, signed images, SBOM, preflight and post-install verification), which turns "runs in our SaaS" into "runs in the customer's account, and they trust it."

Key Responsibilities

Cloud Infrastructure & Kubernetes Operations

Own production infrastructure across fifteen regions on AWS, Azure and GCP: Kubernetes clusters (EKS/AKS/GKE), node pools including GPU nodes for in-platform ML inference, networking, DNS, TLS, CDN/WAF, and cluster-level Helm. Own the data-layer infrastructure our services depend on: MongoDB, Elasticsearch, Redis, RabbitMQ, ClickHouse and object storage; capacity planning, upgrades, backups and restore drills. Operate the observability stack (metrics, logs, traces) as a reliable, cost-controlled service for every engineering team. Own CI runners and build infrastructure; keep pipelines fast, reproducible and secure. Standardise environments through infrastructure-as-code. Every region should be reproducible from a repo, and no change should reach production by hand.

Deployment as a Product (Private Cloud & OEM)

Own the installable Leena package for private-cloud and OEM partner deployments: Terraform modules, Helm charts, signed container images, monthly SBOM regeneration, preflight checks, post-install verification and support bundles. Deliver customer-environment installs (for example, into a customer's own Azure or AWS subscription) with a repeatable runbook and a predictable timeline, and run the upgrade path afterwards. Work with Security, Legal and the partner's engineering team on install requirements, egress allow-lists, air-gap constraints and evidence for customer audits.

Cost & Efficiency

Own a seven-figure annual cloud bill. Baseline it per region and per tenant tier, then reduce it: right-sizing, commitments, storage tiering, egress, and retiring what nobody uses. Report unit economics monthly (cost per region, per tenant, per conversation) with the numbers stated from memory, not from a dashboard you have to open.

Security & Compliance Engineering

Own secrets management (Vault/KMS), image signing, vulnerability remediation SLAs, CIS hardening, network segmentation and least-privilege IAM across all clouds. Provide infrastructure evidence for SOC 2, ISO 27001 and customer VAPT/security reviews; close infrastructure findings on time. Execute the security tooling roadmap (scanning, runtime protection, audit logging) in partnership with the Security function.

Reliability Partnership

Partner with Platform Engineering on reliability: infrastructure-side root cause, capacity headroom, failover and DR testing, and closing postmortem actions that land on infrastructure. Build a DevOps on-call tier that responds fast and pages rarely; aim for a quiet rotation, and treat recurring pages as bugs to be engineered away.

Team Leadership

Lead, mentor and grow the DevOps team; hire the SRE; set clear ownership so every cluster, account, pipeline and bill has a named owner. Replace hero work with runbooks, automation and documentation (Confluence). If a task needs one specific person, that is a risk to remove. Run the roadmap like an engineering team: quarterly goals, a visible backlog, and monthly reporting to the SVP Engineering on reliability, cost and security posture.

Required Skills & Qualifications

Experience: 8–12 years in DevOps, SRE, cloud or infrastructure engineering, including 2+ years managing engineers. Prior experience as a strong IC who still enjoys the keyboard. Kubernetes at depth: Operated multi-cluster, multi-region Kubernetes in production, including upgrades, autoscaling, networking (CNI, ingress, service mesh or equivalent) and stateful workloads. Multi-cloud: Production ownership on at least two of AWS, Azure and GCP; IAM, networking and cost models of each. Infrastructure as code: Expert in Terraform and Helm; disciplined about modules, state, reviews and drift. CI/CD & automation: Built and operated CI/CD at scale (GitHub Actions, GitLab CI, ArgoCD or similar). Fluent in Python, Go or Bash. Observability operations: Run Prometheus/Grafana, ELK/OpenSearch, OpenTelemetry or comparable stacks as services with their own cost and retention decisions. Cost ownership: Has owned and materially reduced a cloud bill, and can walk through how, with numbers. Security fundamentals: Secrets management, image signing and SBOM, vulnerability management, CIS benchmarks, and audit evidence for SOC 2 or ISO 27001. Incident discipline: Has been a responder in production incidents, written postmortems, and driven the fixes to closure. Communication: Explains trade-offs (cost vs. resilience, speed vs. control) plainly to engineers, executives and customers. Education: Bachelor's degree in Computer Science, Engineering, or a related field.

Preferred Qualifications

Shipped software installed into customer-controlled environments: private cloud, on-prem, air-gapped or OEM/partner packaging. Operated GPU infrastructure for ML or LLM inference workloads. Production operation of MongoDB, Elasticsearch, Redis, RabbitMQ or ClickHouse at scale. Multi-tenant B2B SaaS serving large enterprises, with data-residency and regional isolation requirements. Hands-on experience through a compliance audit cycle (SOC 2 Type II, ISO 27001) as the infrastructure owner. FinOps practice: commitment planning, showback/chargeback, unit cost reporting.

How We Interview

We tell candidates this up front because it produces better conversations.

Expect: (1) a deep-dive on an infrastructure estate you personally owned — architecture, what broke, what it cost, and what you changed; (2) a deliberately under-specified design problem from our domain (for example, packaging our platform for a customer's own cloud account), where asking the right questions is scored; (3) a leadership round on how you build a team that runs quietly; (4) a founder round.

Skills: aws,devops,cloud,kubernetes,ci,infrastructure