AI System Engineer
Xpedeon · Mumbai
- Experience3–4 yrs
- SalaryNot disclosed
- Work modeonsite
- Levelmid
- Posted15 Sept 2026
About Xpedeon
Xpedeon is hiring in Mumbai in real estate construction. This role looks for around 3+ years of experience.
Skills
- Python
- retrieval-augmented generation
- SQL
- automated testing
- large language models
- vector search
- embedding-based search
- software engineering
- unit testing
- mocking
- regression testing
- service interfaces
- data access security
The role
A generative AI engineer at a construction and engineering software company designs production agentic systems using retrieval-augmented generation, Python, and SQL, combining language models with validation, business rules, and state management. The role also applies automated testing and cloud deployment to deliver governed natural-language answers over structured business data.
Full job description
About Us: Xpedeon is an integrated purpose-built ERP Suite contributing to digital transformation in the construction and engineering industry across the globe.
Backed by more than two decades of industry experience, Xpedeon has been helping businesses to scale up efficiency, increase profitability, and maximize margins in today’s dynamic business environment. Our integrated solution stack addresses the unique requirements and challenges of building contractors, civil engineering contractors, specialist contractors, and housebuilders by simplifying their business systems with our powerful, comprehensive suite.
More than 30,000 users have gained quantum leaps in productivity and control, improving their operational efficiency and growing their bottom line with Xpedeon.
Job Description: AI Systems Engineer (Chatbots)Team: Product Engineering — AI PlatformReports to: Engineering LeadLevel: Mid–Senior (see notes at bottom)About the roleWe're building an agentic AI system that turns natural-language questions into accurate, governed answers over structured business data. This isn't a thin wrapper around a chat API — it's a multi-stage pipeline with retrieval, guardrails, and session state, and we need someone who can design and build that scaffolding, not just prompt a model well. You'll work on the reliability, safety, and scalability of an AI agent used in production by real customers.What you'll doDesign and build multi-step agent pipelines that combine LLM calls with deterministic logic — retrieval, validation, and business rules — rather than relying on a single prompt to do everything.Implement retrieval-augmented generation (RAG): embedding-based search, ranking, and context assembly that stays within token/latency budgets as the underlying knowledge base grows.Build safeguards around model output — validating, sanitizing, and constraining what the AI is allowed to do or return before it reaches a user or a downstream system.Design for ambiguity: detect when a user's request is underspecified and decide whether to ask a clarifying question, apply a documented default, or refuse — rather than having the system silently guess.Own session and conversation state so multi-turn interactions stay coherent across follow-up questions.Write automated regression tests for AI-driven behavior, including "golden" test sets that catch quality regressions a human might not notice by eyeballing a few responses.Extend the system to new use cases/domains through configuration rather than one-off code, so growth doesn't multiply maintenance burden.Collaborate with frontend, data, and product stakeholders to ground the system in real business rules and correct data access boundaries.Instrument and monitor the pipeline in production — latency, cost, failure modes, and answer quality over time.Required qualifications3+ years of professional software engineering experience, with strong proficiency in Python (or a comparable backend language).Hands-on experience building applications on top of large language models — not just calling a chat API, but designing multi-step logic, structured outputs, and validation around it.Experience with retrieval-augmented generation or vector/embedding-based search.Strong SQL skills and comfort reasoning about data access, correctness, and security boundaries.Solid testing practices — unit testing, mocking external dependencies, and building regression suites for systems with non-deterministic components.Experience designing clean interfaces/contracts between components in a service or pipeline, so systems stay maintainable as they grow.Clear written communication — comfortable documenting assumptions and design decisions rather than guessing silently on ambiguous requirements.Preferred qualificationsExperience with a major LLM provider platform (Google Vertex AI/Gemini, OpenAI, AWS Bedrock, Azure OpenAI, or Anthropic).Experience building a natural-language-to-SQL or natural-language-to-structured-query system specifically.Familiarity with cloud deployment (GCP, AWS, or Azure) and containerization (Docker).Experience with session/cache stores (e.g., Redis) for stateful, multi-turn systems.Frontend familiarity (React/Next.js/TypeScript) sufficient to collaborate across the stack.Domain experience in financial, ERP, or B2B analytics data.What we're explicitly not looking forDeep ML research or model-training expertise — this role is about building robust systems around a foundation model API, not training models.Pure prompt-engineering/no-code AI experience without underlying software engineering skill.Success in the first 90 daysShip at least one end-to-end pipeline component (retrieval, validation, or state management) into production use.Contribute a test suite that catches a real quality or safety regression before it reaches users.Help onboard a new use case/data domain into the system, proving it can extend without a full rewrite.
Seniority notesMid-level: owns individual pipeline components under an established architecture.Senior/lead: additionally owns architectural decisions as the system scales to new domains, sets testing/quality standards, and represents the AI platform in cross-team design discussions.