We are hiring an AI Engineer to build production generative-AI applications for our enterprise clients. You will build RAG pipelines, prompts, LLM integrations and content guardrails as part of a small delivery team under a Solution Architect.
Key Responsibilities
- Build RAG pipelines: chunking, embeddings, hybrid retrieval and custom ranking.
- Write and maintain prompts and structured outputs for multilingual content generation.
- Integrate LLMs behind a routing layer that allows model swaps without code changes.
- Implement content guardrails: prohibited terms, policy rules, content-safety checks and approval workflows.
- Build explainability output — retrieved sources and confidence signals for each generation.
- Run model evaluations and A/B comparisons; track quality, latency and token cost.
- Work with backend, frontend and DevOps engineers; document what you build.
Required Skills
- 4+ years Python; 2–3+ years building LLM applications in production.
- RAG fundamentals: embeddings, vector search, hybrid retrieval, chunking, evaluation.
- Experience with any LLM API (Azure OpenAI, OpenAI, Anthropic, Gemini).
- Experience with any vector database (pgvector, Azure AI Search, Pinecone, Qdrant, Weaviate).
- LangChain or LangGraph (or LlamaIndex / similar).
- FastAPI, PostgreSQL, Docker; Git.
- Any major cloud (Azure preferred).

