Builder 3

ECI Software Solutions · Hyderabad

  • Experience5–20 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Levelsenior
  • Posted18 Sept 2026

About ECI Software Solutions

ECI Software Solutions is hiring in Hyderabad in technology software. This role looks for around 5+ years of experience.

Skills

  • Python
  • SQL
  • TypeScript
  • OpenAI
  • Anthropic
  • PostgreSQL
  • Microsoft SQL Server
  • pgvector
  • Qdrant
  • LangChain
  • LangGraph
  • DSPy
  • dbt
  • Apache Hop
  • Apache Airflow
  • Docker
  • Weights & Biases
  • OpenTelemetry
  • LangSmith
  • Langfuse
  • Braintrust
  • RAG
  • knowledge graphs
  • prompt optimization
  • red-teaming
  • adversarial testing
  • model routing
  • caching
  • CI
  • ERP systems
  • financial data handling
  • agentic architecture
  • guardrail engineering
  • production reliability
  • regression testing
  • human-in-the-loop workflows

The role

A generative AI engineer at an enterprise software company designs agentic architecture and eval frameworks for dependable automation across complex ERP systems, applying guardrail engineering and production reliability practices. The role shapes model routing, tracing, regression evaluation, and human approval paths for financial workflows.

Full job description

BUILDER 3 — Systemic Orchestration

Level: 3 (Lead / Technical Direction)

The Squad: You are the technical lead of the "Modernization Strike Team" — a small squad of Builders rewriting how enterprise software is built, tested and operated using AI. You are not the people manager. You are the person whose judgment the squad borrows when the agent works on Tuesday and fails on Friday.

The Mission: A Builder 2 makes an agent work. You make it work every time, at a cost we can defend, on systems we cannot afford to break. Your job is to turn a collection of clever agents into a dependable production capability across a 20-year-old ERP estate. If you think "it worked in the demo" is a finish line, you will hate it here.

The Builder Levels: A Hierarchy of Acceleration

BUILDER 1 (Professional) — AI-Assisted Execution. Uses AI to write code, tests and documentation 3x faster than a traditional junior. Learning to string together basic chains and prompts.BUILDER 2 (Senior) — Agentic Implementation. The engine room. Doesn't just use AI; builds the agents. Takes a legacy module and builds a LangGraph or CrewAI flow to modernize it with minimal supervision.BUILDER 3 (Lead) — Systemic Orchestration. ? You are here. Designs the Guardrail Layers and the Eval Frameworks. Ensures the agents built by the squad are reliable, cost-effective and scalable.BUILDER 4 (Principal/Architect) — The Meta-Platform. Builds the internal Builder Tools the other three levels use. Works the 12-month horizon of model capability and readies the ERP architecture for autonomous operation.

The honest boundary: B2 owns an agent. B3 owns whether the squad's agents are allowed near production. B4 owns the platform the squad builds on. If you are still measured mainly by the modules you personally shipped, you are a strong B2. If you are measured by whether other people's agents survive contact with production, you are a B3.

The Builder-3 Profile

You are a systems engineer for non-deterministic systems. You are comfortable saying "no, not yet" to a working demo, and you can say exactly what would change your mind.2. The Three Things You OwnWhat You'll Actually DoWhat Good Looks Like

Core DNA (shared foundation, at lead altitude)AI-Native Workflow. You don't just use AI tools fast — you set the squad's bar for what "fast and safe" looks like, and you remove the bottleneck when someone hits one.Agentic Architecture. You still build end-to-end agents yourself, and you do. But your leverage comes from designing the state management, tool interfaces and failure semantics that everyone else builds inside.Legacy-to-Modern Sequencing. You decide which legacy modules get modernized first, based on blast radius, data gravity and how cheaply success can be measured — not on which one looks most fun.Measured Reliability. You define what "reliable enough" means for each workflow: acceptance thresholds, regression suites, human-in-the-loop checkpoints, rollback paths. You know that "100% accurate" is a marketing phrase, and that the engineering version is a measured error rate with a defined containment strategy. The Guardrail Layer Input and output validation, tool-permission scoping, financial and PII data handling, sandboxed execution, human approval gates on irreversible actions, kill switches, blast-radius containment. You assume the agent will eventually do the wrong thing, and you design so that it costs us a log line instead of a General Ledger. The Eval Framework Golden datasets built from real ERP edge cases. Offline evals in CI as a merge gate, online evals in production. LLM-as-judge calibrated against human labels — and you know when that calibration has drifted. Regression detection when a model version, prompt or retrieval index changes underneath you. You move the squad from "prompt vibes" to a number that can fail a build. The Cost & Performance Envelope Model routing, caching, context budgeting, latency budgets. You track token cost per completed business transaction, not per call, and you can tell the CFO why a workflow costs what it costs and where the next 40% comes from. Elective Depth (be deep in one, fluent in all six)Agentic Architect (Systems) — state management and tool-calling interfaces for long-running business processes across complex ERP schemas.UX/Interaction (Product) — generative UIs and browser-use agents; designing where the human stays in the loop and where they don't need to be.Data & Context (Data) — RAG and knowledge-graph architecture: vector stores, chunking strategy, retrieval evals, and the relationship between a Purchase Order and a GL entry.Reliability & Eval (QA) — programmatic prompt optimization (DSPy), red-teaming agents, adversarial testing of financial logic.Infra & Agent-Ops (DevOps) — the factory where agents live: tracing, logging and monitoring the "thoughts" of the AI; deployment and rollback of prompts and models as versioned artifacts.Domain Logic Translation (Business) — turning 50 pages of accounting regulation into structured system prompts and agentic constraints.Set the definition of shippable. Write the standard for when an agent can touch production data, and enforce it in code, not in a wiki page.Build the eval harness and wire it into CI as a gate that can and does block merges.Design the guardrails around regulated and financial logic, including the human approval and rollback paths.Run design reviews. Kill bad agent architectures in week one instead of month four. Be specific about why.Own agent incidents. Pull the traces, find the root cause, and add the eval that makes that failure impossible to reintroduce.Make build-vs-buy calls on frameworks, vendors and models — and revisit them when the ground shifts, which it will.Grow Builder 1s and 2s through code review, design review and pairing. Not through ceremony.

(Targets to be agreed with the hiring manager — these are illustrative, not fixed.)

90 days: Eval harness running in CI on at least two live agentic workflows, with a published reliability baseline. A written, enforced guardrail standard.6 months: Every production agent has a regression suite, a cost-per-transaction number, and a rollback path. Agent-related incidents are down and post-incident evals exist for each.12 months: The squad ships agentic modules on a predictable cadence without you in the critical path of each one.

The Stack

Python, SQL (Postgres/SQL Server), TypeScript, OpenAI/Anthropic, pgvector/Qdrant, LangChain/LangGraph, DSPy, dbt, Apache Hop/Airflow, Docker, Weights & Biases, OpenTelemetry — plus the orchestration layer you'll live in: eval and tracing platforms (LangSmith / Langfuse / Braintrust or equivalent), policy and guardrail tooling, model routing and caching. And any other tool that helps us achieve greatness.

Who You Are

You've been burned. You have shipped an agentic system to real users, watched it fail in a way you didn't predict, and changed your architecture because of it. This is the single strongest signal for this level.Rigor over hype. You love new models and you still ask for the eval numbers first. You know the difference between a cool demo and a production-grade agent handling financial data.Agent-first thinker. You think in capabilities and failure modes, not functions.The manual-task hater. If you do a task twice, you've already started drafting the workflow that kills it.Technically fearless. A 500-table undocumented legacy schema is a puzzle, not a threat.A multiplier. Your best week is one where the squad shipped three things and you wrote one design doc and two brutal, useful reviews.

Why Join Us

The Final Boss of complexity. ERPs are the hardest problem in software for AI. Automate a General Ledger and you've mastered the game.Agentic autonomy. We don't do tickets. We do missions. You choose the models, frameworks and architectures you believe will win.No red tape. If the agent works and the evals pass, we ship. No Change Management Committee — a "Does It Work?" policy. Note that you are the person who defines what "works" means.The Builder Tribe. You will not be the only AI person in the room. You'll lead a squad of people as obsessed with moving 10x faster as you are.