EPAM Systems
1 month ago
Lead GenAI Engineer
Sign up to save this job, get alerts, and apply with an optimized CV.
Company information
- Company
- EPAM Systems
- Location
- Argentina Argentina
- Posted
- 1 month ago
Job description
We are on the lookout for a Lead GenAI Engineer to join our team. In this role, you will oversee the full development, deployment, and operational cycle of enterprise-grade AI-powered applications. The position merges backend engineering, LLM integration, cloud infrastructure, and AI platform operations to deliver scalable GenAI solutions in live production environments. You will collaborate closely with AI/DS, Product, and DevOps teams to build and grow AI-driven applications, upholding reliability, observability, performance optimization, and operational excellence across the complete AI SDLC. The role also involves supporting GenAI-assisted development practices, helping expand the client's enterprise AI SDLC processes, contributing to AI Beauty Chat initiatives through agentic micro-pod delivery models, and carrying out System Steward duties across AI platform initiatives.
Responsibilities
- Design, construct, deploy, and maintain backend services that power AI/LLM-driven applications
- Hold full accountability for GenAI feature delivery, from initial build through to production support
- Connect and manage LLM APIs, including OpenAI, within enterprise production environments
- Develop APIs, orchestration layers, and microservices that enable agentic AI workflows
- Tune LLM systems for latency, resiliency, retries, fallbacks, and cost efficiency
- Set up CI/CD pipelines, observability, monitoring, and logging for AI services
- Work alongside AI/DS, Product, DevOps, and platform teams to simplify delivery and boost reliability
- Operate within Azure cloud environments alongside distributed systems, including Redis, Kafka, and SQL/NoSQL databases
- Facilitate MCP integrations, agentic memory initiatives, and AI orchestration frameworks
- Advocate for GenAI-assisted development practices and support scaling of the client's AI SDLC processes
- Assist with AI Beauty Chat delivery through agentic micro-pod execution models
- Carry out System Steward duties within agentic micro-pods
Requirements
- A minimum of 5 years of experience in a relevant field
- At least one year of experience in a team leadership or management capacity
- Main area of specialization in AI Engineering, with an emphasis on backend systems
- Well-developed background in Python backend development
- Track record of building and running enterprise-grade GenAI/LLM applications from start to finish
- Direct experience working with OpenAI or other LLM APIs in live production settings
- Competence in prompt engineering and various orchestration approaches
- Experience tackling operational hurdles tied to LLMs, including latency, retry logic, fallback handling, observability, and cost control
- Deep familiarity with distributed systems and scalable backend architecture
- Background in CI/CD practices, DevOps tooling, and Azure-based cloud environments
- Experience embedding GenAI into the software development lifecycle, spanning AI-assisted coding, testing, deployment, and release processes
- Practical knowledge of SQL/NoSQL databases, along with Redis and Kafka
- Strong ability to communicate effectively with varied audiences
- Excellent English proficiency (B2 level or higher)
Nice to have
- Exposure to agentic workflow design
- Hands-on background with Databricks and MCP
Required skills
- redis
- kafka
- databricks
- operational excellence
- monitoring
- reliability
- sql
- llm
- prompt engineering
- apis
- devops
- nosql databases
- genai
- cost control
- ci/cd pipelines
- microservices
- observability
- backend services
- cloud infrastructure
- software development lifecycle
- distributed systems
- azure cloud
- openai
- backend engineering
- performance optimization
- devops tooling
- logging
- resiliency
- cost efficiency
- product teams
- llm apis
- latency
- orchestration layers
- ai-assisted testing
- retries
- retry logic
- platform teams
- ai orchestration frameworks
- genai-assisted development
- agentic ai workflows
- mcp integrations
- fallback handling
- fallbacks
- ai-assisted coding
- ai sdlc
- ai platform operations
- scalable genai solutions
- ai beauty chat
- agentic micro-pods
- system steward
- enterprise production environments
- ai/ds
- agentic memory
- ai-assisted deployment
- ai-assisted release processes
- orchestration approaches
- operational hurdles
- scalable backend architecture
- agentic workflow design
Interested in this position?
Create your free account and tailor your CV to match this job.