5+ years commercial Python experience, including 3+ years of hands‑on GenAI/LLM engineering.
Solid engineering fundamentals: building scalable services/APIs (FastAPI, Flask, or Django), writing real test suites (pytest), clean and modular architecture.
Hands‑on experience building RAG pipelines – retrieval, embeddings, vector search.
Working knowledge of agentic patterns: tool‑calling, function‑calling, multi‑step reasoning workflows.
Strong prompt engineering skills, including structured outputs (JSON schemas, Pydantic, Instructor or equivalent).
Experience with at least one major cloud platform (AWS, Azure, or GCP), including its managed AI/ML services (e.g., Bedrock, Azure OpenAI).
An "evals mindset" – you think about relevance, consistency, latency, and cost as real engineering concerns, not afterthoughts.
High‑proficiency written and spoken English – you'll use it daily with clients and teammates.
Nice to have
Experience with orchestration frameworks beyond the basics – LangGraph, LangSmith, LlamaIndex.
Hands‑on with a specific vector database (Pinecone, Weaviate, Milvus, pgvector) beyond "I integrated one once."
Experience building evaluation frameworks or golden‑dataset pipelines specifically (as opposed to just using one).
Exposure to data pipeline work feeding AI systems – understanding how data quality/freshness affects model behavior.
Prior client‑facing / consulting experience in a professional‑services or consulting setup.
Main responsibilities
Design and build Python services and APIs that wrap LLM‑powered functionality.
Build and maintain agentic and RAG pipelines: retrieval, reranking, tool‑calling, multi‑step reasoning, structured outputs.
Care about quality beyond "it works" – evals, observability, and the data that tells you when something regresses.
Work directly with clients: translate fuzzy business requirements into architecture decisions, and explain your technical tradeoffs to non‑technical stakeholders.
Work across a distributed, multi‑market team and client organizations.