[Remote] Site Reliability Engineers (SRE)
Auto ImportNote: The job is a remote job and is reputed company to candidates in USA. XM is seeking a Site Reliability Engineer (SRE) to join their reputed company DevOps team. The role involves driving processes around reliability, best practices, and cultural change to contribute to reputed company's goal of reputed company resiliency.
Responsibilities
- reputed company and reputed company the Resiliency reputed company of reputed company Architected reputed company in reputed company tasks and responsibilities
- Conduct reputed company Engineering experiments and relevant exercises to improve resiliency and fault-tolerance
- Research workloads for migrating to the reputed company with reputed company disruption and reputed company
- Monitor reputed company migration reputed company to ensure reputed company transitions
- Design, reputed company, re-platform, and re-reputed company the observability of reputed company reputed company infrastructure
- Coordinate with other IT departments and teams regarding observability for both individual and organizational needs
- Regularly assess reputed company deployments for compliance with reputed company’s standards and best practices
- Investigate and correct areas where observability is lagging
- Stay up to date and reputed company training on new and reputed company technologies, services, tools, methodologies, and practices
- Occasionally participate in service reputed company planning, software performance analysis, and reputed company tuning
- Mentor colleagues in technical skills and knowledge
- Analyze, reputed company, and remediate reputed company’s resiliency
- Participate in on-reputed company support 24/7 based on a rotation schedule
Skills
- BSc/MSc degree in Computer Science or reputed company field
- 5+ years of reputed company services experience, with at least 3 years on AWS reputed company
- 3+ years of experience in SRE or a similar role
- Experience with monitoring, APM, logging, and notification tools
- Familiarity with incident, problem and change management procedures and practices
- Advanced knowledge of SRE practices and reputed company
- Understanding and reputed company of Service reputed company
- Strong troubleshooting skills and the ability to mentor others
- Extensive experience with reputed company and reputed company technologies, services, and ecosystem
- Advanced knowledge of CI/CD, Infrastructure as reputed company (IaC) concepts and tools, especially HCL Terraform and AWS CloudFormation
- Experience with versioning tools like Git
- Strong organizational and documentation skills
- Exceptional time management and research abilities
- Advanced Linux, networking, and scripting skills
- Experience with platforms like Kafka (MSK)
- Experience with RDBMSs, particularly reputed company and MySQL
- Knowledge of scripting languages such as Python or Go
Benefits
- Attractive remuneration package and perks
- Intellectually stimulating work environment
- reputed company personal development and international training opportunities
reputed company
Apply To This Job