Site Reliability Engineer/L3 Support at SS&C Technologies | New York, United States | Rezi

Site Reliability Engineer/L3 Support at SS&C Technologies

Site Reliability Engineer/L3 Support

SS&C Technologies · New York, United States

Yesterday

Site Reliability Engineer/L3 Support

SS&C Technologies · New York, United States

2 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

This role combines modern Site Reliability Engineering practices with advanced production support responsibilities. You will act as the highest level of operational support (L3), proactively identifying and resolving issues before they impact customers, driving continuous improvement, and working closely with engineering teams to improve the reliability and operability of the platform. This is not a traditional operations role. You will use automation, observability, and engineering best practices to reduce operational toil while helping development teams build resilient, secure services.

Responsibilities

  • Monitor the health, availability, performance, and security of production services.
  • Proactively identify emerging issues using telemetry, logs, metrics, and distributed tracing.
  • Investigate, troubleshoot, and resolve complex production incidents across application and infrastructure layers.
  • Act as the L3 escalation point for operational issues that cannot be resolved by L1 or L2 support.
  • Participate in an on-call rotation for critical production incidents.
  • Lead incident response activities, including coordination, communication, and post-incident reviews.
  • Perform root cause analysis and ensure corrective actions are implemented to prevent recurrence.
  • Develop and maintain operational runbooks, dashboards, alerts, and standard operating procedures.
  • Improve platform observability by enhancing monitoring, alerting, dashboards, and service-level indicators.
  • Work closely with software engineering teams to improve service reliability, scalability, and resilience.
  • Identify opportunities to automate operational tasks and eliminate repetitive manual work.
  • Support production deployments, infrastructure changes, and maintenance activities.
  • Assist with disaster recovery exercises, resilience testing, and operational readiness reviews.
  • Ensure operational activities comply with FedRAMP High security and compliance requirements.
  • Contribute to continuous improvement initiatives across reliability, performance, and operational excellence.

Requirements

  • U.S. Citizenship (required).
  • 3–6 years of experience in Site Reliability Engineering, Production Engineering, DevOps, Platform Engineering, or a senior production support role.
  • Experience supporting mission-critical cloud-based production systems.
  • Strong understanding of Linux operating systems and networking fundamentals.
  • Experience troubleshooting distributed applications running in Kubernetes.
  • Experience with public cloud platforms, preferably AWS.
  • Experience with infrastructure as code and configuration management.
  • Strong scripting or programming skills (e.g. Python, Bash, PowerShell, Go, or similar).
  • Experience using monitoring and observability platforms such as Prometheus, Grafana, CloudWatch, Datadog, Splunk, or OpenTelemetry.
  • Experience analysing application logs, metrics, and traces to diagnose production issues.
  • Understanding of incident management, problem management, and root cause analysis.
  • Strong analytical and troubleshooting skills.
  • Excellent written and verbal communication skills.
  • Experience supporting systems operating under FedRAMP High, DoD IL5/IL6, or similar regulated environments.
  • Experience with Kubernetes in production.
  • Experience with AWS services including EKS, RDS, IAM, CloudWatch, Route 53, VPC networking, and AWS Backup.
  • Experience with CI/CD pipelines and deployment automation.
  • Knowledge of service mesh technologies such as Istio.
  • Familiarity with security best practices including IAM, least privilege, vulnerability management, and compliance monitoring.
  • Experience with PagerDuty, Jira Service Management, or similar incident management platforms.
  • AWS certification (Associate or Professional) is desirable.

Skills

  • Site Reliability Engineering
  • Production Engineering
  • DevOps
  • Platform Engineering
  • Linux operating systems
  • Networking fundamentals
  • Kubernetes
  • AWS
  • Infrastructure as code
  • Configuration management
  • Python
  • Bash
  • PowerShell
  • Go
  • Prometheus
  • Grafana
  • CloudWatch
  • Datadog
  • Splunk
  • OpenTelemetry
  • Incident management
  • Problem management
  • Root cause analysis
  • EKS
  • RDS
  • IAM
  • Route 53
  • VPC networking
  • AWS Backup
  • CI/CD pipelines
  • Istio
  • PagerDuty
  • Jira Service Management

Location

  • REMOTE

Work Type

  • REMOTE
  • Hybrid Work Model

Experience Level

  • 3–6 years of experience

Salary/Compensations

  • 110000 USD to 130000 USD

Benefits

  • 401k Matching Program
  • Professional Development Reimbursement
  • Flexible Personal/Vacation Time Off
  • Sick Leave
  • Paid Holidays
  • Medical
  • Dental
  • Vision
  • Employee Assistance Program
  • Parental Leave
  • Discounts on fitness clubs
  • Discounts on travel

About the Company

  • As a leading financial services and healthcare technology company based on revenue, SS&C is headquartered in Windsor, Connecticut, and has 27,000+ employees in 35 countries. Some 20,000 financial services and healthcare organizations, from the world's largest companies to small and mid-market firms, rely on SS&C for expertise, scale, and technology.

Equal Opportunity

  • SS&C Technologies is an Equal Employment Opportunity employer and does not discriminate against any applicant for employment or employee on the basis of race, color, religious creed, gender, age, marital status, sexual orientation, national origin, disability, veteran status or any other classification protected by applicable discrimination laws.