Senior Generative AI Engineer
About the role
The Role
We are looking for a hands-on Generative AI and LLM Engineer to design, build and deploy production-grade AI applications. The role will focus on developing LLM-powered products, Retrieval-Augmented Generation (RAG) pipelines and intelligent agent workflows using Python, modern AI frameworks and cloud infrastructure.
You will own solutions from requirement understanding and architecture through development, deployment and production monitoring. The ideal candidate combines strong Python backend development experience with practical knowledge of LLMs, vector databases, API integrations, Docker and at least one major cloud platform.
What You Will Do
Design and develop scalable LLM-powered applications using Python.
Build RAG pipelines using document processing, embeddings, vector databases, semantic search and reranking.
Develop AI-agent and multi-agent workflows with tool calling, memory, orchestration and human approval steps.
Integrate LLMs with internal systems, external APIs, databases and enterprise applications.
Evaluate and select suitable foundation models based on accuracy, latency, cost, security and business requirements.
Improve prompt quality, retrieval accuracy, response time and token usage.
Implement safety guardrails, output validation, access controls and fallback mechanisms.
Build automated evaluation frameworks to measure response quality, hallucination, relevance and reliability.
Containerize applications using Docker and deploy them on AWS, Azure or GCP.
Implement monitoring and LLMOps practices for model performance, cost, latency, errors and production usage.
Collaborate with product, engineering and business teams to convert requirements into reliable AI solutions.
Document technical architecture, design decisions, APIs and operational processes.
What Success Looks Like
Production-ready AI applications that are accurate, secure and maintainable.
RAG systems that retrieve relevant information and reduce hallucinations.
AI-agent workflows that reliably complete business tasks and integrate with existing systems.
Measurable improvements in response quality, latency and inference cost.
Clear monitoring of application performance, usage, errors and model behaviour.
What We're Looking For
Bachelor's or Master's degree in Computer Science, Engineering or a related discipline, or equivalent practical software-development experience.
6-9 years of professional software-development experience, including strong hands-on experience with Python.
Experience developing backend services and integrating REST APIs.
Hands-on experience building LLM or Generative AI applications.
Practical experience implementing RAG using embeddings, semantic search and vector databases.
Experience with LangChain, LangGraph, LlamaIndex or a comparable LLM application framework.
Experience integrating foundation models through APIs such as OpenAI, Anthropic Claude, Gemini or Azure OpenAI.
Working knowledge of vector databases such as Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS or pgvector.
Experience with Docker and deployment on at least one cloud platform: AWS, Azure or GCP.
Understanding of prompt engineering, hallucination reduction, output validation and LLM evaluation.
Strong understanding of software engineering practices, Git, testing, debugging and clean code.
Ability to communicate technical solutions clearly to both technical and non-technical stakeholders.
Technical Frameworks and Toolkit
Experience building agentic or multi-agent workflows using LangGraph, AutoGen, CrewAI, Semantic Kernel or similar frameworks.
Experience with Hugging Face Transformers, PyTorch or fine-tuning techniques such as LoRA or QLoRA.
Knowledge of model serving and inference frameworks such as vLLM, TGI or Ollama.
Experience with Kubernetes, CI/CD pipelines and infrastructure automation.
Familiarity with observability or LLMOps tools such as LangSmith, Langfuse, Arize Phoenix, MLflow or Weights & Biases.
Knowledge of reranking, hybrid search, chunking strategies and retrieval evaluation.
Experience implementing AI guardrails, PII protection, prompt-injection prevention and responsible AI practices.
Perks And Benefits of Working With Us
Unlimited PTO.
Please ask us about our very generous parental leave, much above industry standards!
Entrepreneurial culture where pushing limits and taking risks is everyday business.
Open communication with management and company leadership.
Small, dynamic teams = massive impact.
Medical, Dental and Vision coverage for employees.
Access to Disability & Life insurance.
Mental health and wellbeing support.
Annual bonus program.
Employer Stock Purchase Program (ESPP).
Yearly team building experiences.
Mentorship and sponsorship opportunities.
Manager resources and support.
Cogniify is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or any other protected characteristic
Responsibilities
- Design and develop scalable LLM-powered applications using Python
- Build RAG pipelines using document processing, embeddings, vector databases, semantic search and reranking
- Develop AI-agent and multi-agent workflows with tool calling, memory, orchestration and human approval steps
- Integrate LLMs with internal systems, external APIs, databases and enterprise applications
- Evaluate and select suitable foundation models based on accuracy, latency, cost, security and business requirements
- Implement safety guardrails, output validation, access controls and fallback mechanisms
- Build automated evaluation frameworks to measure response quality, hallucination, relevance and reliability
- Containerize applications using Docker and deploy them on AWS, Azure or GCP
Qualifications
- Bachelor's or Master's degree in Computer Science, Engineering or a related discipline, or equivalent practical software-development experience
- 6-9 years of professional software-development experience, including strong hands-on experience with Python
- Experience developing backend services and integrating REST APIs
- Hands-on experience building LLM or Generative AI applications
- Practical experience implementing RAG using embeddings, semantic search and vector databases
- Experience with LangChain, LangGraph, LlamaIndex or a comparable LLM application framework
- Experience integrating foundation models through APIs such as OpenAI, Anthropic Claude, Gemini or Azure OpenAI
- Working knowledge of vector databases such as Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS or pgvector
Benefits
- Unlimited PTO
- Generous parental leave, much above industry standards
- Entrepreneurial culture where pushing limits and taking risks is everyday business
- Open communication with management and company leadership
- Small, dynamic teams = massive impact
- Medical, Dental and Vision coverage for employees
- Access to Disability & Life insurance
- Mental health and wellbeing support
- Annual bonus program
- Employer Stock Purchase Program (ESPP)
- Yearly team building experiences
- Mentorship and sponsorship opportunities
Skills mentioned
About Cogniify
Cogniify helps enterprises move from AI pilots to industrialized impact. Most organizations don’t struggle with ideas or models—they struggle with scaling AI safely, economically, and operationally. AI works in demos, but stalls when it meets real data, real costs, real operations, and real regulation. That’s the gap we solve. We work with leaders across Retail & CPG, Banking & Fintech, Insurance, Manufacturing, Energy & Utilities, Healthcare & Life Sciences, Telecom, Tech & Media, and Supply Chain to design and build the control layers that make AI production-ready. Our work spans: - Data industrialization and trusted enterprise foundations - Real-time sensing and predictive intelligence - Agentic planning and autonomous operations - AI lifecycle management, MLOps & FinOps - Responsible AI, governance, and regulatory readiness We don’t sell tools. We don’t run pilots for the sake of pilots. We help organizations: - scale AI without cloud bill shock - move from dashboards to decisions - Turn insights into execution - Pass audits before regulators ask - and make AI accountable at the P&L level Cogniify exists to close the value gap between experimentation and execution so AI becomes a durable capability, not a recurring initiative. From pilots to production. From insight to action. From AI ambition to AI control.