LATAM | Senior AI Engineer
Location: Latin America (ideally São Paulo, Bogotá, Mexico City, Buenos Aires or Chile)
Employment Type: Full-time
Salary: Up to $130,000 USD/year
About the role
We're working with a fast-growing, venture-backed healthcare AI company whose AI agents already run in production inside clinical workflows across Latin America.
They are looking for a Senior AI Engineer to take these agents to the next level. This is not an LLM integration role. The core of the job is training and specializing models, structuring clinical knowledge, and building evaluation systems that go beyond standard metrics.
This is a highly technical, hands-on Individual Contributor role on a small team. There is no people management.
What you'll do
- Fine-tune and distill small LLMs to specialize and accelerate the agent stack, owning the full cycle: data curation, training, evaluation, and deployment
- Design clinical ontologies and structured knowledge representations, so decision support becomes deterministic and auditable
- Decide when a problem needs better context, model training, or deterministic code, and justify the call on quality, latency, and cost
- Build router models and task decomposition that make the easy 80% cheap and fast and the hard 20% correct
- Push evaluation beyond the standard playbook: adversarial and counterfactual evals, calibrated LLM-as-judge, failure taxonomies, and uncertainty estimation
- Apply context engineering as measured engineering (retrieval, compression, memory, structure), not prompt tweaking
- Own these systems in production (monitoring, cost, failure modes) and set the bar for how agents are built and evaluated
Must-have requirements (all required)
-
You have personally fine-tuned, trained, or distilled a model end to end, and the result was used in production or in a real product. You handled the data, the training, and the evaluation yourself.
- Not sufficient on its own: calling LLM APIs, prompt engineering, running a managed fine-tuning job without owning the data and evaluation, or AI training / RLHF annotation work for AI labs.
-
You have built and shipped AI agents or LLM systems that real users rely on in production, not only prototypes or internal demos.
-
You have designed an evaluation system yourself and can describe a specific failure it caught that standard metrics missed.
-
Hands-on experience with RAG and knowledge graphs or structured knowledge (e.g., ontologies, Neo4j, graph-based retrieval) used to ground agents.
-
Strong Python as your primary language for AI/ML work, including frameworks such as PyTorch or Hugging Face.
-
3+ years of hands-on AI/ML engineering, not counting experience only integrating LLM APIs into applications.
- Fluent English, used in your daily work
- Based in Latin America
Nice to have
- Healthcare or clinical AI experience (e.g., clinical data, SNOMED CT, FHIR, ICD)
- Experience building router models or tiered model architectures
- Ability to read research papers and judge what holds up in production
- Early-stage startup experience
- Daily use of AI coding tools (Claude Code, Cursor, Codex)
This role is not for you if:
- Your AI experience is mainly integrating LLM APIs (OpenAI, Anthropic, Gemini) into products, with no model training
- Your fine-tuning experience is limited to proofs of concept or tutorials
- Your background is mainly backend, full-stack, or infrastructure, and AI is a recent addition
- You are looking for a leadership or people-management role