AI & HPC Infrastructure Engineer
| Hours | Full-time, Part-time |
|---|---|
| Location | United, WV United, West Virginia open_in_new |
About this job
Job Description
Senior AI & HPC Infrastructure Engineer Remote (US) | $150,000 - $170,000 Base + Bonus
\nOur client is a highly respected global consulting and research organization that supports leading commercial enterprises, public sector institutions, and professional services firms. As part of a significant investment in AI infrastructure, they are expanding their internal high performance computing (HPC) and GPU capabilities to support next-generation machine learning and large language model (LLM) initiatives.
\nThis is an exciting opportunity to join a small, highly skilled infrastructure team at a pivotal stage of growth. The successful candidate will play a key role in scaling a GPU environment from 8 to 32 NVIDIA H200 GPUs while helping shape the organization's long-term AI and HPC strategy.
\nThe role is heavily project-focused, with approximately 85% dedicated to engineering, architecture, and platform development activities, and a smaller proportion supporting operational needs.
\nThe Opportunity
\nYou'll work at the intersection of traditional HPC, research computing, and modern AI infrastructure, supporting both analytical workloads and large-scale model training environments.
\nKey responsibilities include:
\n- \n
- Designing, deploying, and maintaining GPU-accelerated computing infrastructure \n
- Supporting large-scale AI/ML and LLM training environments \n
- Managing Linux-based HPC clusters and associated services \n
- Administering and optimizing parallel file systems, particularly IBM Spectrum Scale (GPFS) \n
- Managing NVIDIA GPU platforms, CUDA, cuDNN, NCCL, and related tooling \n
- Supporting resource scheduling through SLURM and related technologies \n
- Performance tuning for distributed and multi-GPU workloads \n
- Building automation, monitoring, reporting, and operational tooling \n
- Collaborating with researchers, data scientists, and technical stakeholders to translate business requirements into infrastructure solutions \n
- Evaluating emerging AI infrastructure technologies and recommending future platform enhancements \n
Required Experience
\nWe're particularly interested in candidates who combine deep infrastructure expertise with an understanding of how end-users consume HPC resources.
\nEssential Skills
\n- \n
- Strong Linux systems administration background \n
- Experience supporting HPC or research computing environments \n
- Hands-on NVIDIA GPU infrastructure experience \n
- Strong experience with GPFS / IBM Spectrum Scale \n
- Familiarity with job schedulers such as SLURM, LSF, or similar \n
- Experience supporting distributed compute environments \n
- Ability to lead projects independently with minimal oversight \n
- Strong troubleshooting experience across compute, storage, networking, and hardware layers \n
- Excellent communication skills with the ability to explain complex technical concepts to non-technical stakeholders \n
Highly Desirable
\n- \n
- Experience supporting AI/ML infrastructure or LLM platforms \n
- Kubernetes and container orchestration experience \n
- MLOps tooling exposure (MLflow, Kubeflow, etc.) \n
- Experience tuning large model training and inference workloads \n
- Bright Cluster Manager \n
- Ansible \n
- Docker, Apptainer, or Singularity \n
- SAS platform experience alongside GPU/HPC expertise \n
Ideal Background
\nThis role is particularly well suited to someone who:
\n- \n
- Has approximately 5-10 years of relevant infrastructure experience \n
- Is currently operating at a strong mid-level and ready for a senior step forward \n
- Can own and deliver significant technical projects independently \n
- Comes from an HPC, research computing, higher education, government, scientific computing, or technical enterprise environment \n
- Enjoys balancing infrastructure engineering with emerging AI technologies \n
Team & Culture
\nYou'll join a close-knit team of experienced infrastructure specialists covering HPC operations, automation, applications, and platform engineering. The group operates in a highly collaborative remote-first model, with team members distributed across the United States.
\nA notable area of planned growth is AI and LLM infrastructure expertise, making this a high-visibility hire with significant opportunity for progression and influence.
\nIf you're an HPC infrastructure engineer looking to move into a highly visible AI-focused environment while remaining close to the hardware, architecture, and platform engineering side of the stack, we'd love to hear from you.