Staff Engineer - Distributed Systems

Gohighlevel · India

  • Experience10–11 yrs
  • SalaryNot disclosed
  • Work moderemote
  • Posted25 Sept 2026

About Gohighlevel

Gohighlevel is hiring in India in technology software. This role looks for around 10+ years of experience.

Skills

  • Node.js
  • Go
  • TypeScript
  • Kubernetes
  • Google Cloud Platform
  • Google Kubernetes Engine
  • Google Cloud Pub/Sub
  • Google Cloud Tasks
  • MongoDB
  • Firestore
  • ClickHouse
  • Elasticsearch
  • Redis
  • SQL
  • NoSQL
  • distributed systems
  • queueing architectures
  • async architectures
  • consistency models
  • idempotency

The role

A distributed systems engineer at an AI-powered business software company designs resilient backend services and APIs using Node.js and Kubernetes, shaping queueing architectures and storage systems for high-throughput communication infrastructure. The role also applies Redis and Google Cloud Platform to strengthen failure handling and scalability across critical production systems.

Full job description

The Role

About HighLevel:HighLevel is an AI-powered business operating system that gives agencies, entrepreneurs and SMBs the infrastructure to build, automate and scale. Today, HighLevel supports SMBs across 150+ countries, fueling community-driven growth rooted in real customer outcomes.

To date, businesses operating on HighLevel have generated over $7 billion in ecosystem value, demonstrating the impact of shared infrastructure at scale. By centralizing conversations, automation and intelligence into one system, we help businesses move faster, reduce complexity and execute efficiently.

Behind the platform, HighLevel powers more than 4 billion API hits and 2.5 billion message events daily. With 250 terabytes of distributed data, 250+ microservices and over 1 million domain names supported, our architecture is built for performance, resilience and long-term scalability.

Our people

With over 2,000 team members across 10+ countries, HighLevel operates as a global, remote-first organization built for speed and ownership. We value initiative, clarity and execution, creating space for ambitious people to build systems that support millions of businesses worldwide. Here, innovation thrives, ideas are celebrated and people come first, no matter where they call home.

Our impact

Every month, HighLevel enables more than 1.5 billion messages, 200 million leads and 20 million conversations for the more than 1 million businesses we support. Behind those numbers are real people building independence, expanding opportunity and creating measurable impact. We’re proud to be a part of that.

Learn more about us on our YouTube Channel or Blog Posts

Role Summary:

We’re hiring a Senior Full Stack Engineer to design, build, and scale the backend systems and frontend interfaces that power the Threads & Composer experience inside HighLevel Conversations. You’ll spend most of your time on backend and API design, data flows, and performance — while still owning the Vue 3 UI layer end-to-end. Your work will touch critical paths that process billions of messages per month, demanding reliability, throughput, and correctness at scale.

Team & System Overview:

The Conversations organization powers multi-channel communication across SMS, Email, WhatsApp, and DMs — handling over 2B messages each month. The Threads & Composer team owns the inbox and composer surfaces, and the backend services that deliver and store those interactions.You’ll work across Node.js microservices, GCP-based infra (GKE, Pub/Sub, Cloud Tasks), and MongoDB/Firestore storage layers, ensuring every message and UI state stays consistent, fast, and recoverable.

Responsibilities:

Own the architecture health of a billion-scale distributed system, its failure modes, capacity limits, consistency guarantees, and the interactions between 50+ deployments that no single team can see

Approve critical-path designs. Changes that touch the system's core go through you, not as bureaucracy, but as the person accountable for the whole staying sound. When there's a disagreement, you make your case on merit

Hunt gaps proactively, single points of failure, unbounded queues, missing idempotency, thundering herds, quiet data-loss windows, and drive the fixes before they become incidents

Build the parts nobody else can. No sprint tickets. You prototype the risky architectural bets yourself, ship the remediation after serious incidents, and pair into the gnarliest cross-team bugs, roughly a quarter to a third of your time in code, all of it on the hardest problems

Make resilience a property of the system, not a heroic act, degradation strategies, backpressure, isolation boundaries, capacity models that survive 10%+ month-over-month growth

Raise the teams around you. Design reviews that teach, post-mortems that change architecture (not just add alerts), and patterns that 80+ engineers build on

Set the standard for how AI-assisted engineering works safely on systems this critical, where a bad merge doesn't cost a demo, it costs real businesses their revenue

The terrain:

Runtime: Node.js (TypeScript), Go, thousands of pods on GKE

Messaging & async: GCP Pub/Sub, Cloud Tasks, Redis

Storage: MongoDB, Firestore, ClickHouse, ElasticSearch

Scale: 21.5B automation actions/month, 2.6B messages/month, 28.5K req/s peaks, ~226B async events across the org

Requirements:

10+ years of engineering experience with deep, hands-on ownership of large-scale distributed systems, comparable scale strongly preferred: hundreds of services or thousands of instances, billions of daily events

You've carried sole accountability for a production system through real failures, not adjacent to it, not advising on it. Owned it

Deep command of queueing and async architectures, delivery semantics, ordering, backpressure, idempotency, exactly-once myths and at-least-once realities

Strong with multiple storage engines (SQL and NoSQL), you reason about consistency models, indexing at scale, and when each engine is the wrong choice

Expert-level depth in Redis or comparable in-memory systems, including their failure modes under memory pressure and network partition

Production experience on Kubernetes at scale, resource limits, autoscaling behavior, what actually happens when a node pool dies

Exceptional design communication, docs, diagrams, and RCAs that drive decisions across multiple teams

Fluent in Node.js and/or Go, enough to prototype your own proposals and ship fixes on the critical path

What Success Looks Like

The critical paths of Workflows have named owners, capacity models, and tested failure modes, because you made it so

Incident count trends down while traffic grows double-digit percent month over month

Your design reviews are the ones engineers want their proposals to survive

Your scope has expanded on results, more systems, more surface, more trust

What you're built for

You've been injured in production. A lot. You've owned distributed systems at serious scale, through the outages, the migrations, the 3 AM discoveries, and every scar changed how you design

You've operated at staff scope, whatever your title said, the engineer everyone routed the hardest systems questions to

You think in failure modes by default: when you see a design, you instinctively ask what happens at the tail, under partition, at 10x load

You write design docs and RCAs that people reference years later, clear trade-offs, honest risks, real recommendations

You can disagree with a team and still make them better, influence through rigor and respect, not title

You'd rather prevent ten incidents quietly than be the hero of one loudly

Bonus Points

You've made AI agents genuinely productive on complex systems, and know how to keep AI-generated code from becoming AI-generated incidents

GCP-native experience: Pub/Sub, Cloud Tasks, GKE, Firestore

You've done this job before under another name, "the systems person," principal engineer, architect-who-still-codes

Experience taking a 0→1 system to production alongside hardening mature ones

Additional Information

#LI-Remote #LI-HB1

EEO Statement:

The company is an Equal Opportunity Employer. As an employer subject to affirmative action regulations, we invite you to voluntarily provide the following demographic information. This information is used solely for compliance with government recordkeeping, reporting, and other legal requirements. Providing this information is voluntary and refusal to do so will not affect your application status. This data will be kept separate from your application and will not be used in the hiring decision.

We encourage you to review our Privacy Policy before submitting your application