https://bayt.page.link/vYp9BaDsgeBXhZZb6
Create a job alert for similar positions

Job Description

Introduction
At IBM, work is more than a job – it’s a calling: To build. To design. To code. To consult. To think along with clients and sell. To make markets. To invent. To collaborate. Not just to do something better, but to attempt things you’ve never thought possible. Are you ready to lead in this new era of technology and solve some of the world’s most challenging problems? If so, lets talk.

Your Role and Responsibilities We are seeking an experienced Site Reliability Engineer (SRE) to join our team. The ideal candidate will have 5 to 8 years of experience in ensuring the reliability, availability, and performance of critical services and systems. This role involves building and maintaining infrastructure, automating processes, and responding to incidents to ensure our systems run smoothly and efficiently.
Key Responsibilities:
Infrastructure Management:

  • Design, build, and maintain scalable, resilient infrastructure using cloud platforms (AWS and Azure).
  • Manage and optimize Kubernetes clusters, containers, and microservices.
  • Implement Infrastructure as Code (IaC) using tools like Terraform (Must), Ansible (Good to have), or CloudFormation(Good to have)
Automation & CI/CD:
  • Maintain automated CI/CD pipelines to ensure rapid, safe, and reliable delivery of software.
  • Automate repetitive tasks, processes, and workflows to increase efficiency and reduce human error.
  • Implement and maintain monitoring, logging, and alerting systems to ensure visibility into system performance.
Cost Optimization:
  • Set up monitoring and reporting tools to track cloud spending in real-time.
  • Regularly review the architecture and operations to identify areas where costs can be reduced. This includes evaluating new tools, services, or practices that could lead to further cost savings.
  • Collaborate with development teams to ensure that cost-efficient practices are followed in software design and deployment.
  • Recommend and manage the purchase of reserved instances, savings plans, or other discounts offered by cloud providers to reduce costs for long-term workloads.
Incident Response & Troubleshooting:
  • Respond to and resolve incidents in a timely manner, ensuring minimal downtime and impact on customers.
  • Perform root cause analysis and post-mortem reviews to prevent recurrence of issues.
  • Collaborate with development teams to improve system reliability through proactive issue identification and resolution.
Performance Optimization:
  • Monitor system performance and capacity, and implement improvements to optimize efficiency and scalability.
  • Analyze and improve application performance, ensuring high availability and low latency.
  • Security & Compliance:
  • Ensure security best practices are followed across the infrastructure.
  • Implement security controls and monitoring to protect against vulnerabilities and threats.
  • Work with compliance teams to ensure systems adhere to regulatory requirements.
Collaboration & Communication:
  • Work closely with software engineers, product managers, platform team, Global Support and other stakeholders to ensure system reliability aligns with business goals.
  • Provide guidance and mentorship to junior SREs and other team members.
  • Document processes, procedures, and best practices for the broader team.


Required Technical and Professional Expertise


  • Looking for a Java backend developer with an experience of 5-7 years of industry experience
  • Strong in Java, Data Structures & Algorithms, OOPS.
  • Experience in webservice(SOAP, REST) and cloud concepts
  • Should have good communication skills, ability to understand the business needs and must have sound analytical skills
  • Experience in Docker, Kubernetes will be an added advantage.
  • Strong proficiency in version control systems, such as Git.


Preferred Technical and Professional Expertise


  • Analytical mindset with the ability to troubleshoot complex issues and provide effective solutions.
  • Experience working in Agile/Scrum development environments.
  • Proactive in staying updated with industry trends, technologies, and best practices.
  • Write and maintain unit and integration tests to ensure the reliability of code.
  • Collaborate with quality assurance teams to identify and resolve bugs and issues.
  • Work closely with product managers, designers, and other developers to understand project requirements and deliver high-quality products

Job Details

Job Location
Bengaluru India
Company Industry
Other Business Support Services
Company Type
Employer (Private Sector)
Employment Type
Unspecified
Monthly Salary Range
Unspecified
Number of Vacancies
Unspecified
You have reached your limit of 15 Job Alerts. To create a new Job Alert, delete one of your existing Job Alerts first.
Similar jobs alert created successfully. You can manage alerts in settings.
Similar jobs alert disabled successfully. You can manage alerts in settings.