logo

AI Job Summary

Evaluate and characterize next-generation GPU and AI accelerator platforms for Oracle Cloud Infrastructure, working hands-on in the lab to bring up hardware, enable AI software stacks, and analyze system performance. Requires 8+ years developing software infrastructure for large-scale AI systems and deep expertise in GPU/CPU architecture, CUDA, ROCm, and debugging hardware-software integration issues. You'll profile AI workloads, identify performance bottlenecks, and provide architectural recommendations across platforms. US-based, paying $96,800–$306,400 annually.

Written from this posting by Neural Jobs AI. The full description is below.

AI Resume Tailoring Sign in to use this AI Cover Letter Sign in to use this

Job Description

Oracle Hardware Platform Development Engineering is seeking a highly driven AI Systems Engineer to evaluate and characterize next-generation GPU and AI accelerator platforms for Oracle Cloud Infrastructure (OCI). This is a hands-on engineering role focused on bringing up new hardware platforms, enabling AI training and inference software stacks, running representative workloads, and analyzing system performance under real operating conditions.

The engineer will identify whether workloads are HBM/memory-bandwidth, compute, scale-up, or scale-out bound, while characterizing power, thermals, memory behavior, utilization, scaling, and performance efficiency. Working directly in the lab, you will debug hardware/software integration issues, design and execute experiments, and develop data-driven insights that explain system behavior beyond benchmark results.

A key part of the role is comparative architecture analysis across GPUs and emerging AI accelerators. You will evaluate architectural tradeoffs and translate performance findings into clear, actionable recommendations on which platforms are best suited for specific AI training and inference workloads. You will work closely with internal hardware and software teams as well as technology partners to help shape Oracle’s next generation of high-performance AI infrastructure.

Position Overview:

This position is ideal for someone who loves deep systems engineering, debugging complex hardware–software interactions, and optimizing performance at every layer of the ML stack. You will play a pivotal role in enabling the training and deployment of next-generation LLMs and generative AI models.

Responsibilities

Required Qualifications 

  • Solid knowledge of AI / GPU or/and AI/CPU platform architecture and their capabilities. 
  • Experience with the architecture, design, and implementation of modern server platforms consisting of multiple architectures and vendors, including x86 and ARM server architectures.
  • Strong communications skills and ability to clearly communicate complex technical issue across engineering disciplines as well as clearly and succinctly articulate issues for executives. 
  • Experience and understanding of the latest high-speed busses and interconnect used in modern Compute and AI platforms. Familiarity with their startup connectivity and operational robustness as well as performance metrics.

  • Debugging & Reliability: Troubleshoot complex hardware–software interaction issues, including vLLM compilation failures on ROCm, CUDA memory leaks, distributed runtime failures, and kernel-level inconsistencies.

  • Profiling & Performance Analysis: Conduct detailed profiling of compilation graphs, training workloads, and runtime execution to optimize performance and eliminate bottlenecks.

 

Preferred Qualifications 

  • Minimum of 8+ years of experience in developing software infrastructure for large scale AI systems.

  • Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience).

  • Strong debugging skills and experience in analyzing and triaging AI applications from the application level to the hardware level.

  • Hands-on experience maintaining or building ML training stacks involving CUDA, ROCm, NCCL, XLA, or similar technologies.

  • Experience in benchmarking AI workloads across different architectures.

  • Background in working with the large scale clusters

  • Good understanding on DL frameworks internal PyTorch, TensorFlow, JAX, and Ray

 

 

 

Qualifications

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $96,800 to $306,400 per annum. May be eligible for bonus, equity, and compensation deferral.


Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - IC5


Do you match this job?

Here is what this employer asked for. Sign in and we will fill in your half.

  • Role AI Systems Engineer (OCI/AI Infrastructure)
  • Experience 3-4 years
  • Education Bachelor Degree, or equivalent experience
  • Work type On-site
  • Location United States
Check my match (free)
Oracle
Enterprise Software · 500+ Members · Austin, TX, United States

Oracle is a global enterprise-software and cloud-computing company.

Oracle develops databases, cloud infrastructure, business applications, developer tools, analytics, and industry-specific technology platforms.

Job Overview

Approx. salary range

189K – 287K

Our estimate — this employer did not publish a salary

Our estimate, not the employer’s. Worked out from the middle half of 51 comparable roles on Neural Jobs that did publish a salary, in the same field, country and experience band. The real figure for this job may be different.

Eligibility
United States Right to work in the United States required.
Workplace
On-site
Job Posted:
2 days ago
Job Type
Full Time
Education
Bachelor Degree, or equivalent experience
Experience
3-4 years

Share This Job: