Software Engineer II - AI Infrastructure

Black Eagle Defense
Fort Meade, Maryland, United StatesFull-time$143,000–$200,000Posted Aug 28, 2026

About the role

Job Description SALARY RANGE $143,000 - $200,000/year DUTIES As a successful candidate for the Software Engineer II - AI Infrastructure role, you will support the development, operation, and evolution of the next generation of AI infrastructure that enables innovation across the customer organization. As part of a full-stack engineering team, you will design, implement, and maintain scalable platform capabilities that serve as the foundation for AI-powered applications and services. Your efforts will focus on AI inference infrastructure while supporting a broader ecosystem that includes advanced analytics, retrieval-augmented generation (RAG), autonomous agents, and emerging AI technologies. In this role, you will independently design, develop, deploy, and optimize infrastructure components that deliver reliable, secure, and high-performance AI capabilities at scale. You will collaborate with engineers, platform teams, and stakeholders to enhance platform reliability, drive adoption of modern technologies and engineering practices, and ensure AI services remain scalable, observable, and operationally resilient. Through cloud engineering, automation, systems integration, and platform development, you will help deliver the infrastructure that powers mission-critical AI solutions across the enterprise. Required Skills SKILLS * Design, implement, and optimize infrastructure supporting AI model inference at scale

Develop, deploy, and maintain production AI services and applications, including retrieval-augmented generation (RAG), autonomous agents, and emerging AI technologies

Analyze ambiguous requirements and define scalable, maintainable solutions for complex systems and operational challenges

Drive the adoption of modern technologies, engineering standards, and best practices across development teams

Implement monitoring, logging, and observability capabilities to improve visibility into AI platform performance and reliability

Automate infrastructure provisioning, deployment, and configuration management using Infrastructure-as-Code principles

Ensure the availability, reliability, scalability, and performance of AI platform components and supporting services

Contribute to the implementation of security best practices for AI systems, services, and data environments

Design and integrate platform capabilities that support enterprise AI initiatives and operational requirements

Collaborate with engineers, platform teams, and stakeholders to improve AI infrastructure and service delivery

Troubleshoot complex infrastructure, platform, and application issues within production environments

Provide technical guidance, knowledge sharing, and informal mentorship to junior engineers

Support the continuous improvement and modernization of AI infrastructure, cloud environments, and platform operations

Contribute to the full lifecycle of AI platform development, from design and implementation through deployment and sustainment QUALIFICATIONS Eight (8) years of experience as a SWE in programs and contracts of similar scope, type, and complexity are required. A Bachelor's degree in Computer Science or a related discipline from an accredited college or university is required. Four (4) years of additional SWE experience on projects with similar software processes may be substituted for a bachelor's degree. Additional requirements: * Proven experience building, deploying, and maintaining production systems at scale

Experience designing and optimizing high-volume web application architectures for performance, scalability, and reliability

Strong background in systems integration across diverse technologies, platforms, and services

Hands-on experience with cloud engineering and solution deployment within AWS environments

Proficiency in administering and deploying applications within Kubernetes-based environments

Strong Python development skills for automation, infrastructure, and application development efforts

Experience implementing observability and monitoring solutions using technologies such as APM, OpenTelemetry, Grafana, and Prometheus

Familiarity with CI/CD pipelines, automation frameworks, and DevOps best practices

Strong understanding of infrastructure automation, deployment strategies, and operational excellence principles

Strong change management, stakeholder engagement, and organizational influence skills

Ability to operate effectively within ambiguous environments and establish structure for evolving requirements

Strong analytical, troubleshooting, and problem-solving skills

Excellent written and verbal communication skills

Ability to collaborate effectively across multidisciplinary engineering and operational teams

Experience supporting the full lifecycle of cloud-native applications and platform services from design through production operations Desired Skills NICE-TO-HAVES * Experience with AI inference serving technologies such as vLLM, LiteLLM, or similar platforms

Experience developing solutions using agentic AI frameworks such as LangChain or comparable technologies

Knowledge of vector databases, embedding models, and semantic search architectures

Experience designing and supporting retrieval-augmented generation (RAG) solutions and AI-enabled applications

Familiarity with large language model deployment, optimization, and inference workflows

Experience with high-performance computing environments and distributed systems architectures

Knowledge of scalable data processing, distributed computing, and resource optimization techniques

Experience supporting enterprise AI platforms and machine learning infrastructure

Familiarity with emerging AI technologies, frameworks, and platform capabilities

Experience integrating AI services and infrastructure into cloud-native environments and production systems

Responsibilities

  • Support the development, operation, and evolution of AI infrastructure
  • Design, implement, and maintain scalable platform capabilities
  • Independently design, develop, deploy, and optimize infrastructure components
  • Collaborate with engineers, platform teams, and stakeholders
  • Implement monitoring, logging, and observability capabilities
  • Automate infrastructure provisioning, deployment, and configuration management
  • Contribute to the implementation of security best practices
  • Troubleshoot complex infrastructure, platform, and application issues

Qualifications

  • Eight years of experience as a Software Engineer
  • Bachelor's degree in Computer Science or related discipline
  • Proven experience building, deploying, and maintaining production systems
  • Experience designing and optimizing high-volume web application architectures
  • Strong background in systems integration across diverse technologies
  • Hands-on experience with cloud engineering and AWS environments
  • Proficiency in administering applications within Kubernetes-based environments
  • Strong Python development skills

Skills mentioned

PythonAWSKubernetesInfrastructure as CodeCI/CDOpenTelemetryPrometheusGrafanaRetrieval-Augmented GenerationModel Serving

About Black Eagle Defense

Black Eagle Defense is a mission-first, Cybersecurity-focused, Information Technology organization. Our goal is to provide premier technical service, targeted staffing, and consulting to our public and private sector clients to ensure operational success.

IT Services and IT Consulting11-50 employeesAnnapolis, MD