BUILDER 2 QA Engineer: Agentic QA and Reliability Engineer
ECI Software Solutions · Hyderabad
- Experience5–8 yrs
- SalaryNot disclosed
- Work modeonsite
- Posted29 Sept 2026
About ECI Software Solutions
ECI Software Solutions is hiring in Hyderabad in technology software. This role looks for around 5+ years of experience.
Skills
- Python
- pytest
- Playwright
- SQL
- TypeScript
- OpenAI
- Anthropic
- LangChain
- LangGraph
- DSPy
- pgvector
- Qdrant
- Docker
- CI/CD
- OpenTelemetry
- mutation testing
The role
A QA engineer at an enterprise software company builds Python agents for agentic test generation, ERP workflow validation, and test stabilization, using Playwright and pytest to execute reliable coverage. The role also applies mutation testing and OpenTelemetry to investigate flakiness and production failures.
Full job description
BUILDER 2 — Agentic Implementation (QA)
Level: 2 (Senior / High-Velocity Execution)
Track: Quality & Reliability Engineering
The Squad: You are one of a few Builders in a “Modernization Strike Team” rewriting how enterprise software is built, tested and operated using AI. Small squad, no layers, short conversations.
The Mission: Build the agents that turn a ticket into a working, executed, stable test — against ERP products where the business logic is 20 years old and nobody wrote it down. You are not writing test cases. You are building the machinery that writes, runs and maintains them. If you want a role where you receive a build and click through it, you’re going to hate it here.
The Builder Levels
Builder is not a title, it’s a measure of autonomy, architectural complexity and agentic leverage. There is one ladder — Dev, QA, Data and DevOps are electives on it, not separate tracks.
BUILDER 1 (Professional) — AI-Assisted Execution. Uses AI to write code, tests and documentation far faster than a traditional junior. Learning to string together chains and prompts.BUILDER 2 (Senior) — Agentic Implementation. ? You are here. The engine room. Doesn’t just use AI; builds the agents. Takes a product area and builds the agentic pipeline that generates, runs and stabilizes its tests with minimal supervision.BUILDER 3 (Lead) — Systemic Orchestration. Designs the guardrail layers and eval frameworks that decide whether the squad’s agents are allowed near production.BUILDER 4 (Principal/Architect) — The Meta-Platform. Builds the internal Builder tools the other three levels use.
The honest boundary: a traditional Senior QA Engineer is measured by tests written and defects found. A Builder 2 is measured by the pipeline they built — and by what stopped escaping to production because of it. You write agents, in Python, that other engineers depend on.
The Builder-2 (QA) Profile
You are an engineer who specialises in proving things work. The distinction between “dev” and “QA” work is not one we maintain: you read the application code, you write production-grade Python, and you build agents. The difference is what you point them at.2. What You Own
Core DNAAI-Native Workflow. You use AI to accelerate everything you do as a default, not an experiment. You have opinions about which tools earn their place.Agentic Implementation. You build, debug and deploy autonomous agents — LLM-driven test generation, Playwright browser-use agents, self-healing locator strategies — against complex ERP screens and workflows.Legacy-to-Modern Bridge. Some of what you’ll test is a modern web rewrite. Some is a desktop application older than your career. Both are puzzles.Evidence Over Confidence. A generated test that has never been executed is not a test. You know your own false-positive and flake rates rather than claiming there aren’t any.
The ticket-to-test pipeline. An agent that reads a requirement or defect, understands the affected area, and produces a real test — grounded in the actual application, not in what the model imagines the application looks like. Hallucinated locators are the failure mode that kills these systems; preventing them is your problem.
Execution and stabilization. Generation is the easy half. Your pipeline isn’t done when a test file exists or a ticket gets created — it’s done when the test has run against a real build, passed repeatedly, and earned its place in the suite. Flake is the enemy, because a suite nobody trusts is worse than no suite.
Coverage that means something. You can tell the difference between a suite that covers the code and a suite that covers the risk. You use mutation testing and real defect history rather than counting tests, and you argue for the edge cases that actually bite: fiscal calendars, multi-currency, partial shipments, period-end adjustments.
Where You Might Lean InTest Generation & Grounding — the retrieval and context problem. How does an agent know what the screen really contains, what the business rule really is, and what changed in this release?Adversarial & Red-Team — you build agents whose job is to break things. Programmatic prompt optimization (DSPy), fuzzing business logic, hunting the input nobody thought of.Test Infrastructure & Agent-Ops — the factory the test agents run in: parallel execution, environment and data management, tracing, and the observability to know why a run went sideways.Legacy Surface Automation — desktop and thick-client applications where modern tooling doesn’t reach and the automation has to be built rather than installed.
What You’ll Actually Do
Build test-generation agents that take a ticket and produce an executable test grounded in the live application.Close the loop. Make generation, execution, stabilization and reporting one pipeline rather than four disconnected steps with a human carrying work between them.Automate the untestable. Wrap legacy business logic — the modules everyone has avoided because there’s no spec — in characterization tests that capture what the system actually does today.Kill flake. Find the root cause rather than adding a retry. Retries are how suites die.Red-team the financial logic. Build the agents that go looking for the case where the ledger doesn’t balance.Fix your own agents in production. Read the traces, find why it went sideways, make it not happen again.
What Good Looks Like
(Illustrative — real targets agreed with the hiring manager.)
30 days: Shipped your first agent-generated test into a running suite — generated, executed, stable.90 days: Owning the ticket-to-test pipeline for a product area end to end, with a measured flake rate and a real coverage picture.6 months: Other Builders reuse what you built. Escaped defects in your area are measurably down, and you can show why.
The Stack
Python, pytest, Playwright, SQL (Postgres/SQL Server), TypeScript, OpenAI/Anthropic, LangChain/LangGraph, DSPy, pgvector/Qdrant, Docker, CI/CD pipelines, OpenTelemetry, eval and tracing platforms (LangSmith / Langfuse / Braintrust or equivalent), desktop automation frameworks where the application demands it. And any other tool that helps us achieve greatness.
Who You Are
You’ve shipped one. You have built at least one autonomous agent or complex LLM workflow that real people used for real work — not a weekend demo. You know how it failed, because it did.You write real code. Python that goes to production and gets reviewed like anyone else’s. If your automation experience stops at a record-and-playback tool, this isn’t the role.The manual-task hater. You have a physical reaction to repetitive work. If you do it twice, you’ve already started drafting the automation.Technically fearless. A 500-table undocumented legacy schema isn’t complexity, it’s a puzzle about to get solved. You aren’t intimidated by spaghetti code — you’re the one who brings the sauce.Rigor over hype. You love the new models and you still ask for the numbers. You know the difference between a cool demo and a pipeline people rely on every release.Suspicious by instinct. Your first question about any green build is what it didn’t check.
Why Join Us
Quality is engineering here, not a stage gate. You are a Builder on the same ladder, with the same levels and the same expectations as everyone else in the squad. You build; you don’t sign off on other people’s work.The Final Boss of complexity. ERPs are the ultimate challenge for AI. If you can build agents that validate a General Ledger, you’ve solved the hardest version of this problem.Agentic autonomy. We don’t do tickets. We do missions. You’ll have real freedom over the models, frameworks and agent architectures you think will win.No red tape. If the agent works and the evals pass, we ship. No Change Management Committee — a “Does It Work?” policy.The Builder Tribe. You won’t be the only AI person in the room. You’ll be part of a small squad equally obsessed with moving faster, where your automation hacks are your biggest contribution rather than a hidden side project.