Principal AI Engineer
About the role
Principal AI Engineer
NYC (Hybrid 3 Days)
Full-Time
Seeking a Staff/Principal AI Engineer to lead enterprise-scale Agentic AI solutions for Fortune 500 clients. Must have deep hands-on experience building and deploying production-grade LLM applications, autonomous agents, RAG systems, MCP servers, vector databases, and distributed backend systems.
Roles & Responsibilities
Design and build agentic systems: Lead the architecture and implementation of tool-calling agents that combine retrieval, structured reasoning, and secure action execution with least-privilege access.
Productionize LLM applications: Build retrieval pipelines, prompt synthesis, response validation, and self-correction loops, backed by rigorous evaluation.
Own the full stack: Deliver the data pipelines, backend services, distributed compute, and orchestration layer that agentic systems depend on — not only the model invocation.
Engineer for reliability and governance: Build validator models, adversarial test suites, and policy checks; enforce deterministic fallbacks and rollback strategies; instrument continuous evaluation.
Optimize for cost and latency: Drive measurable improvements in token efficiency, response time, and unit economics against defined SLOs.
Codebase ownership: Build, maintain, and review high-quality Python and SQL, with an emphasis on reusable components, scalability, and performance.
Cloud integration: Deploy AI applications on AWS, Azure, or GCP with optimized resource usage and robust CI/CD.
Cross-functional collaboration: Partner with product owners, data scientists, and business SMEs to define requirements and deliver impactful AI products.
Mentoring and technical leadership: Set engineering standards and share knowledge across the team, raising the bar on AI and software engineering practice.
Required :
8–14 years of software engineering experience, with strong hands-on large-scale Python
Working depth in at least one systems or backend language — Go, Rust, Java, or C/C++ — and the judgment to know when to reach for it
Strong data structures and algorithms.
Strong understanding of APIs, microservices, and system design
Hands-on experience building and operating data pipelines and production-grade distributed systems.
Agentic AI and LLMs
2+ years of hands-on LLM engineering, with at least couple agentic system you designed and took to production
Production experience with agent frameworks — LangGraph, Google ADK, CrewAI, Claude Agent SDK, or equivalent — and the fluency to move between them as the ecosystem evolves
Experience building MCP (Model Context Protocol) servers and tool-calling interfaces
RAG from first principles: chunking strategy, embeddings, vector and hybrid retrieval, reranking, and response validation
Strong experience with vector databases (Milvus, Pinecone, Weaviate, FAISS, etc. or cloud equivalents)
Design of guardrails and reliability patterns — validators, policy checks, self-correction loops, deterministic fallbacks, circuit breakers, and rollback paths
Optimization
Deep familiarity with token optimization and context-window management — context shaping, pruning, and compaction
Latency and cost optimization through caching, model routing, batching, streaming, and parallel tool calls
Performance testing and tuning systems against defined SLOs
Evaluation
Experience building evaluation frameworks for LLM systems — offline eval sets, continuous online evaluation, and regression detection
Instrumentation and traceability suitable for regulated enterprise environments using tools like LangSmith, Langfuse, etc.
Responsibilities
- Lead the architecture and implementation of tool-calling agents
- Build retrieval pipelines, prompt synthesis, response validation, and self-correction loops
- Deliver the data pipelines, backend services, distributed compute, and orchestration layer
- Build validator models, adversarial test suites, and policy checks
- Drive measurable improvements in token efficiency, response time, and unit economics
- Build, maintain, and review high-quality Python and SQL code
- Deploy AI applications on AWS, Azure, or GCP
- Partner with product owners, data scientists, and business SMEs
Qualifications
- 8–14 years of software engineering experience
- Strong hands-on large-scale Python experience
- Working depth in at least one systems or backend language
- Strong data structures and algorithms knowledge
- Strong understanding of APIs, microservices, and system design
- Hands-on experience building and operating data pipelines
- 2+ years of hands-on LLM engineering experience
- Production experience with agent frameworks
Skills mentioned
About IMR Soft LLC
IMR Soft is an outcomes-driven technology services and strategic resourcing partner helping enterprises build, modernize, operate, and scale critical technology capabilities. Headquartered in Princeton, New Jersey, with strong delivery depth across the U.S. and India, we support clients across BFSI, pharmaceuticals and life sciences, manufacturing, automotive, telecom, and technology-led industries. Our capabilities span managed services, SAP transformation, AI-led product engineering, cloud and infrastructure, data and analytics, application support, QA automation, DevOps, SRE, and strategic technology resourcing. We help clients move from capacity challenges to accountable delivery models — including specialist teams, project pods, 90-day pilots, managed-services engagements, and long-term transformation support. What makes IMR Soft different is our ability to combine strategic talent access with delivery governance, senior accountability, and measurable outcomes. We do not simply provide resources; we help clients improve operational resilience, reduce backlog, accelerate delivery, strengthen data trust, modernize platforms, and scale execution with confidence. We Engineer Outcomes — by converting technology capability, delivery discipline, and strategic talent into measurable business value.