E

Principal Systems Engineer

Eeho.Fa.Us2.Oraclecloud.Com|Jobsearch·See all Eeho.Fa.Us2.Oraclecloud.Com|Jobsearch jobs on Jobr

Posted 1 day ago

OfficeNashville, TN, United StatesSE80k - 187k USD

We are seeking a technical operations leader to join the AI Infra Operations team as a Principal Systems Engineer supporting GPU infrastructure in Oracle Cloud Infrastructure (OCI). In this role, you will lead the development and maintenance of automation and operational tooling for GPU fleets across multiple regions, ensuring high availability. You will collaborate closely with engineering and operations teams to build robust automation, observability, and reliability solutions while continuously improving GPU operations.

  • Build and maintain automation and operational tooling for OCI GPU infrastructure across multiple geographic regions.

  • Drive collaboration with software engineers, hardware teams, and operations partners to maintain a highly available GPU fleet.

  • Build and improve monitoring, alerting, and diagnostics for GPU fleet health, performance, capacity, and utilization using tools such as Grafana.

  • Serve as the senior escalation point for complex GPU host and repair issues

  • Participate in incident response and root-cause analysis to remove blockers affecting GPU capacity, availability, and regional deployments.

  • Continuously improve AI2 Ops processes, GPU fleet automation, and OCI region build readiness.

  • Participate in on-call rotations and provide support for critical infrastructure issues.

  • Document operational procedures, automation workflows, troubleshooting guides, and runbooks.

  • Build and improve AI agents, ensuring safe rollout, execution and monitoring.

  • Mentor and guide junior engineers in operational best practices, provide senior technical support, and drive constant improvement.

 

Required Qualifications:  

 

  • 8+ years of software operations or infrastructure automation experience with strong proficiency in Python and Bash.

  • Expert Linux administration experience, particularly Ubuntu and Oracle Linux, in large-scale production environments.

  • Strong understanding of distributed systems, including peer-to-peer, node-to-node, and service-to-service communication patterns.

  • Strong data-center and host-lifecycle experience, including provisioning, validation, repair workflows, hardware replacement, and fleet recovery.

  • Strong problem-solving and troubleshooting skills.

  • Excellent communication and teamwork skills.

  • Experience with observability tooling, including metrics, logging, dashboards, and alerting.

  • Experience with AI agents and tooling

  • Experience leading on-call operations and incident response.

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

 

Preferred Skills:

 

  • Hands-on experience with GPU infrastructure, including NVIDIA and AMD based systems, and familiarity with GPU shapes or instance families.

  • Experience operating or automating GPU, compute, or other large-scale cloud infrastructure.

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $79,900 to $187,000 per annum. May be eligible for bonus and equity.


Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
As part of Oracle's onboarding process and consistent with applicable law, US-based employees are required to complete identity verification, which involves the collection and processing of their biometric information. Accommodations to this requirement may be granted following an individualized assessment.
E
Eeho.Fa.Us2.Oraclecloud.Com|Jobsearch
View company page
Apply smarter with Jobr

Jobr aggregates jobs directly from company career portals — no middlemen. Our team applies on your behalf with AI-tailored resumes, reviewed by a human before submission.

Direct from company career pages
AI-personalised cover letters
Human review before every submit
Application tracking & follow-ups