Builds production ML/AI services end to end: multimodal ETL/ELT pipelines, RAG systems, and CV/NLP/LLM model integration via vLLM, SGLang, PyTorch and ONNX with GPU inference optimization, plus async Python services, prompt engineering, guardrails, testing and CI/CD. Remote or hybrid from Moscow.