Senior Software Engineer – GPU Infrastructure Automation

GTN Technical Staffing
Dallas, Texas, United StatesFull-timePosted Aug 27, 2026

About the role

Senior Software Engineer – Fleet Automation / GPU Infrastructure

Location: Dallas, TX

Work Model: Hybrid, 3 days onsite

Relocation: Available

Compensation: base + bonus

Benefits: 100% company-paid benefits

Employment Type: Direct Hire

Overview

Our client is seeking a Senior Software Engineer to join its Fleet Automation team supporting large-scale HPC and GPU infrastructure.

This team builds the software, automation, and internal platforms used to provision, configure, monitor, and manage hundreds of high-performance GPU and CPU compute nodes. The environment sits at the intersection of software engineering, infrastructure, and distributed systems, with a strong focus on eliminating manual operational work through scalable automation.

This is a hands-on engineering role for someone who enjoys building backend services while also understanding how those services interact with Linux systems, physical hardware, networking, storage, and GPU infrastructure.

What You’ll Do

Design and build fleet automation platforms for provisioning, configuration, validation, and lifecycle management of GPU and CPU compute nodes.

Develop internal services and APIs that automate hardware deployment, imaging, remediation, and decommissioning.

Build reliable backend services using Go, C#, and/or TypeScript.

Design data models and persistent state for automation workflows using relational and NoSQL databases.

Develop and maintain CI/CD pipelines for infrastructure and configuration changes.

Automate hardware validation and testing across large-scale compute environments.

Build observability, monitoring, dashboards, and alerting using tools such as Prometheus, Grafana, Alertmanager, and ELK.

Partner closely with Infrastructure, Network, Operations, and Research teams to identify operational pain points and automate repeatable processes.

Participate in on-call rotations and support incident response, root-cause analysis, and post-incident reliability improvements.

Identify systemic infrastructure issues and develop engineering solutions that improve fleet reliability, efficiency, and scalability.

Required Experience

5+ years of software engineering experience building production backend services, infrastructure platforms, or automation tooling.

Strong development experience with at least one of the following:

Go

C#

TypeScript

Experience designing APIs, backend services, and distributed or stateful systems.

Strong experience with relational and/or NoSQL databases.

Solid Linux systems knowledge, including:

Networking

Storage

Process management

System troubleshooting

Ubuntu and/or RHEL environments

Experience building and maintaining CI/CD pipelines.

Hands-on experience with production monitoring and observability platforms such as:

Prometheus

Grafana

Alertmanager

ELK

Strong troubleshooting and problem-solving skills across both software and infrastructure environments.

Preferred Experience

Experience supporting GPU, HPC, AI/ML, or large-scale compute infrastructure.

Familiarity with NVIDIA technologies such as:

DCGM

nvidia-smi

NVIDIA Container Toolkit

Experience with bare-metal provisioning and hardware lifecycle automation.

Exposure to event-driven architectures and messaging platforms such as Kafka.

Experience working with infrastructure, network, SRE, or platform engineering teams.

Bachelor’s degree in Computer Science, Software Engineering, or equivalent practical experience.

Why Consider This Opportunity?

Work directly on large-scale GPU and HPC infrastructure supporting advanced AI workloads.

Build automation platforms that have direct impact on infrastructure reliability and scalability.

Highly technical environment combining software engineering, systems engineering, and infrastructure automation.

100% company-paid benefits.

Relocation assistance available for candidates moving to Dallas.

Hybrid schedule with three days per week onsite.

Responsibilities

  • Design and build fleet automation platforms for provisioning, configuration, validation, and lifecycle management of GPU and CPU compute nodes
  • Develop internal services and APIs that automate hardware deployment, imaging, remediation, and decommissioning
  • Build reliable backend services using Go, C#, and/or TypeScript
  • Design data models and persistent state for automation workflows using relational and NoSQL databases
  • Develop and maintain CI/CD pipelines for infrastructure and configuration changes
  • Automate hardware validation and testing across large-scale compute environments
  • Build observability, monitoring, dashboards, and alerting using tools such as Prometheus, Grafana, Alertmanager, and ELK
  • Partner closely with Infrastructure, Network, Operations, and Research teams to identify operational pain points and automate repeatable processes

Qualifications

  • 5+ years of software engineering experience building production backend services, infrastructure platforms, or automation tooling
  • Strong development experience with at least one of the following: Go, C#, TypeScript
  • Experience designing APIs, backend services, and distributed or stateful systems
  • Strong experience with relational and/or NoSQL databases
  • Solid Linux systems knowledge, including networking, storage, process management, and system troubleshooting
  • Experience building and maintaining CI/CD pipelines
  • Hands-on experience with production monitoring and observability platforms such as Prometheus, Grafana, Alertmanager, and ELK
  • Strong troubleshooting and problem-solving skills across both software and infrastructure environments

Benefits

  • 100% company-paid benefits
  • Relocation assistance available for candidates moving to Dallas
  • Hybrid schedule with three days per week onsite

Skills mentioned

GoBackend DevelopmentDistributed SystemsSQLNoSQLLinuxCI/CDAutomationPrometheusGrafana

About GTN Technical Staffing

Global adaptive solutions for IT, engineering, and managed services. Connecting top talent with leading companies. Offices in Phoenix, Dallas, Houston.

Staffing and Recruiting201-500 employeesPhoenix Dallas & Houston, Texas