Automation Engineer

Roche · Hyderabad

  • Experience5–8 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Levelsenior
  • Posted2 Sept 2026

About Roche

Roche is hiring in Hyderabad in pharma biotech. This role looks for around 5+ years of experience.

Skills

  • Python
  • FastAPI
  • asyncio
  • REST APIs
  • GraphQL APIs
  • Git
  • Docker
  • Kubernetes
  • OpenTelemetry
  • LangChain
  • LangGraph
  • LlamaIndex
  • CrewAI
  • AutoGen
  • DSPy
  • FAISS
  • Milvus
  • Qdrant
  • Pinecone
  • Weaviate
  • Pgvector
  • vLLM
  • Hugging Face TGI
  • Ollama
  • NeMo Guardrails
  • Guardrails AI
  • Pydantic
  • Instructor
  • Ragas
  • DeepEval
  • Langfuse
  • LangSmith
  • Phoenix
  • Unstructured
  • Apache Tika
  • LlamaParse
  • PDFPlumber
  • Model Context Protocol
  • Agent2Agent communication standards
  • Amazon Bedrock
  • Amazon SageMaker
  • Amazon OpenSearch Service
  • AWS Step Functions

The role

A generative AI engineer at a pharmaceutical and clinical research company designs agentic workflows using retrieval-augmented generation, LangChain, and Amazon Bedrock, optimizing autonomous systems for regulated document automation. The role develops secure LLM applications with vector databases and clinical trial standards.

Full job description

Responsibilities:

Agentic Architecture: Design, deploy, and scale multi-agent orchestration systems and autonomous workflows using cutting-edge frameworks.

Advanced RAG Pipelines: Build and optimize advanced retrieval-augmented generation (RAG) pipelines over massive, heterogeneous datasets (structured and unstructured).

State and Memory Management: Implement robust state management, short/long-term memory systems, and self-correction/reflection loops within agent networks.

Evaluation and Guardrails: Create and implement robust evaluation metrics, observability pipelines, and guardrails for content quality, hallucination reduction, bias mitigation, and safety standards.

Performance Optimization: Monitor and optimize AI inference cost, latency, throughput, token usage, and overall system reliability.

Security and Access Control: Implement robust access controls, data encryption, user authentication, and prompt injection mitigation across all LLM workflows.

Collaboration and Best Practices: Document and share reusable agent patterns, prompt libraries, and engineering components across cross-functional technical teams.

Requirements:

Demonstrated experience taking ownership of ambiguous tasks, successfully driving small to medium initiatives, and acting as a technical mentor.

Proven track record of engaging in knowledge-sharing initiatives, speaking at internal technical events, and navigating group dynamics in diversified settings.

Technical Skills:

GenAI Development: Advanced prompt engineering, fine-tuning, RAG/GraphRAG, schema-constrained outputs, function/tool-calling, and semantic caching.

Agentic Frameworks: LangChain, LangGraph, LlamaIndex, CrewAI, AutoGen, and DSPy.

Vector Databases: FAISS, Milvus, Qdrant, Pinecone, Weaviate, and Pgvector.

LLM Serving and Infra: vLLM, Hugging Face TGI, and Ollama (for local development).

Guardrails and Validation: NeMo Guardrails, Guardrails AI, Pydantic, and Instructor.

LLM Ops and Observability: Ragas, DeepEval, Langfuse, LangSmith, and Phoenix.

Data Parsing: Unstructured, Apache Tika, LlamaParse, and PDFPlumber.

Programming and DevOps: Python (FastAPI, asyncio), REST/GraphQL APIs, Git, CI/CD pipelines, Docker, Kubernetes, and OpenTelemetry.

Emerging Protocols: Model Context Protocol (MCP) and Agent2Agent communication standards.

AWS Ecosystem (Baseline Experience):

Amazon Bedrock: Foundation model access, custom configurations, and managed agent workflows.

Amazon SageMaker: Fine-tuning, hosting, evaluating, and deploying open-source LLMs.

Amazon OpenSearch Service: Vector search, hybrid search, and enterprise retrieval infrastructure.

AWS Step Functions: Multi-step orchestration and state machine management for hybrid AI/traditional pipelines.

Additional Qualifications:

Problem-Solving: An analytical mindset capable of breaking down highly abstract, ambiguous logic loops into predictable agent behaviors.

Collaboration: Ability to thrive in a fast-paced, product-focused agile engineering environment alongside data scientists and product owners.

Communication: Strong technical writing and communication skills for documenting complex system architectures and cross-functional collaboration.

Bonus: Domain and Regulatory Knowledge:

Experience or strong familiarity with the following clinical trial standards will give you a significant advantage:

Clinical Document Automation: Automating the full clinical document generation workflow, specifically translating protocols to Clinical Study Reports (Protocol CSR).

Data Lineage Workflows: Experience transforming raw clinical data and statistical outputs into structured regulatory documents (SDTM/ADaM/ARD TLG CSR).

CDISC Standards: Deep understanding of CDASH (CRFs), SDTM, ADaM, ARD/ARM, and Define-XML.

Regulatory Submissions: Familiarity with ICH M11 (protocol/SoA), ICH E3 (CSR), and eCTD Module 5

Compliance Frameworks: Designing software to align strictly with GxP, 21 CFR Part 11 and ICH guidelines, ensuring total auditability, traceability, and explainability.