Network Operations Centre Engineer - Alert and Incident Management (Chennai)

Athenahealth Technology Private · Chennai

  • Experience3–7 yrs
  • SalaryNot disclosed
  • Work modeonsite
  • Posted25 Sept 2026

About Athenahealth Technology Private

Athenahealth Technology Private is hiring in Chennai in healthcare. This role looks for around 3+ years of experience.

Skills

  • Linux administration
  • system administration
  • Oracle
  • MySQL
  • PostgreSQL
  • fault management
  • incident management
  • Infrastructure-as-Code
  • Agile

The role

A network operations engineer at a healthcare technology company monitors production systems, applies Linux administration and manages incident management for cloud infrastructure. The role also uses Infrastructure-as-Code and fault management.

Full job description

NOC Engineer (Alert & Incident Management) Join us as we work to create a thriving ecosystem that delivers accessible, high-quality, and sustainable healthcare for all.

Position Summary: We are looking for an NOC Engineer to join our Alert & Incident Management team , part of Cloud Infrastructure Engineering (Network Operations Centre) . You will promote a teaching and learning culture within the team and serve as a NOC liaison to all internal stakeholders . As a NOC liaison, you will identify and execute opportunities to adopt recent processes and improve operational practices within the team.

This is NOT a directly client-facing role . However, every action you take will be in the interest of delivering an amazing user experience for athenahealth clients and internal users alike .

About the Team: NOC (Alert & Incident Management) is a Global Operations team responsible for 24/7/365 support of the core athenaNet application (Collector & Clinicals) , all supporting Platform and Product technology services , the entire supporting infrastructure across public and private cloud for core athenaNet, athenahealth Enterprise applications , and more services as the environment continues to grow. The team is focused on availability, reliability, stability, scalability, customer experience , and solving large-scale distributed computing problems , while maintaining and improving the performance of the applications and infrastructure it supports. The team ideally spends approximately 50% of its time reacting to issues and 50% of its time making the overall system better through continuous improvement, automation, process enhancement, and operational excellence.

The team operates within an agile environment and focuses on maintaining and evangelizing Infrastructure-as-Code , while continuously improving availability, scalability, reliability, and customer experience.

Essential Job Responsibilities Alert

Management & Resolution 40%: Monitor, identify, triage, investigate, and resolve alerts and operational issues in a timely manner. Perform incident management activities to restore services and minimize impact to applications, infrastructure, and users. Support 24/7/365 operational requirements for core athenaNet, supporting Platform and Product technology services, infrastructure, and Enterprise applications.

Apply strong system administration and Linux administration skills to troubleshoot and resolve production issues. Work with database technologies such as Oracle, MySQL, PostgreSQL, and similar technologies when troubleshooting operational issues.

Use Fault

Management and Monitoring tools to identify, analyze, and respond to alerts and system events. Support applications and infrastructure in a production environment, maintaining availability, reliability, stability, and scalability.

Incident

Management 30%: Coordinate incident response, troubleshooting, escalation, communication, and resolution across relevant teams. Work effectively with cross-functional groups and teams to achieve common operational and business goals. Serve as a NOC liaison to internal stakeholders, ensuring effective communication and coordination during operational events. Deliver clear communications to both technical and non-technical stakeholders.

Continuous

Improvement 20%: Identify opportunities to improve processes, operational practices, system reliability, and team effectiveness. Promote a teaching and learning culture within the team. Identify and execute opportunities to adopt new processes and improve existing operational workflows.

Contribute to Infrastructure-as-Code practices and the continuous improvement of the supported environment. Create and maintain technical documentation and Standard Operating Procedures (SOPs). Turnover, Team Meetings, Vendor Management & Ticket Work 10%: Complete operational handovers, participate in team meetings, manage vendor-related activities, and perform ticket-related work.

Follow established operational best practices and contribute to maintaining consistent and reliable operational processes.

Additional Job Responsibilities

Participate in standard weekday shifts and rotational weekend shifts as required. Provide operational support outside of the regular schedule when business or operational needs arise. Maintain effective shift turnover and handoff to ensure continuity of operations.

Work with internal teams and vendors to coordinate issue resolution and operational activities. Support the adoption of new tools, processes, and technologies that improve operational efficiency. Contribute to knowledge sharing and continuous learning within the NOC team.

Assist with maintaining accurate tickets, operational records, documentation, and incident information. Support initiatives focused on availability, reliability, stability, scalability, and customer experience. Help improve the overall system by identifying recurring issues and opportunities for operational .