Job Description

Solve complex problems related to infrastructure cloud services and build automation to prevent problem recurrence. Design, write, and deploy software to improve the availability, scalability, and efficiency of Oracle products and services. Design and develop designs, architectures, standards, and methods for large-scale distributed systems. Facilitate service capacity planning and demand forecasting, software performance analysis, and system tuning.
 

At Oracle Cloud Infrastructure (OCI), we build the future of the cloud for Enterprises as a diverse team of fellow creators and inventors. We act with the speed and attitude of a start-up, with the scale and customer-focus of the leading enterprise software company in the world. Compute is one of the core organisations within OCI. We are responsible for providing Compute power . VMs and BMs. Cloud pretty much cannot exists without our org. The Compute org comprises of a family of critical foundational infrastructure services that drive OCI’s hardware lifecycle activities

Work with Site Reliability Engineering (SRE) team on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services. Responsible for the design and delivery of the mission critical stack, with focus on security, resiliency, scale, and performance. Authority for end-to-end performance and operability. Partner with development teams in defining and implementing improvements in service architecture. Articulate technical characteristics of services and technology areas and guide Development Teams to engineer and add premier capabilities to the Oracle Cloud service portfolio. Understand and communicate the scale, capacity, security, performance attributes, and requirements of the service and technology stack. Demonstrate clear understanding of automation and orchestration principles. Act as ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs). Utilize a deep understanding of service topology and their dependencies required to troubleshoot issues and define mitigations. Understand and explain the affect of product architecture decisions on distributed systems. Professional curiosity and a desire to a develop deep understanding of services and technologies.

Responsibilities include but not limited to
Incident Management
Support and troubleshooting of Staging/Production environments
Response and Resolve incidents as per SLA's
Organise, Anticipate, Plan and work as On-Call in shifts for multiple services (Open to work in shifts & shows flexibility)
Maintain Service High Availability
Release Management
Test and Deploy solutions and automate to replace manual processes
Build and maintain deployment tools/procedures
Zero downtime deployments and a high availability mindset
Define and build innovative solution methodologies and assets around infrastructure, cloud migration and deployment operations at scale.
Work with service teams to resolve complex issues that require troubleshooting and knowledge of code.
Keep documentation up to date and resolving similar tickets with lower turnaround time and within SLA
Ensure production security posture
Ensure monitoring is robust and effective
Change Management
Perform Root Cause Analysis

Qualifications:
  • Bachelors in computer science and Engineering or related engineering fields
  • 6+ years of experience delivering and operating large scale, highly available distributed systems.
  • 5+ years of experience with Linux System Engineering
  • 4+ years of experience with Python/Java building infrastructure Automations
  • Understanding of Networking, Cloud Computing, Load Balancers
  • Strong Infrastructure troubleshooting skills
  • Experience in CICD, Cloud Computing and networking
  • Hands on experience at Monitoring/Instrumentation tools (Prometheus/Grafana etc)
  • Career Level - IC4

    Apply for this Position

    Ready to join ? Click the button below to submit your application.

    Submit Application