SENIOR PYTHON / AI AGENT ENGINEER
HeartStamp · Las Vegas ·
- Seniority
- Senior
- Employment
- Contract
- Category
- AI engineering
- Experience
- 2+ years
HeartStamp · Las Vegas ·
pwc · Kolkata
MGNY Consulting Corp. · Europe
practice-by-numbers · Gurugram, IND
Action1 · Spain
Compensation: $24k – $42k • No equity
HeartStamp | Remote | Full-time contract | Start immediately
Own the agent that sits at the center of a live consumer product.
HeartStamp is an AI-native platform that helps people create one-of-a-kind greeting cards, professionally printed and mailed to the door, or delivered as an animated digital card. We launched in the US in 2026 and we are shipping fast: digital 3D cards, direct mail, invitations, and a mobile app.
We are built for Millennials and Gen Z who find traditional cards generic and emotionally hollow. A customer talks to Stampy, our conversational assistant, and comes away with a card that looks like the one they would have made if they had the time, the talent, and the taste.
We are a small, high-output distributed team. Every engineer here owns a real surface and ships to real customers.
Stampy is not a chatbot bolted onto a storefront. It is a stateful, tool-using agent that captures what a customer means, resolves it against a live catalogue of tens of thousands of cards, drives image generation, and hands off into a print-ready pipeline. It is the product.
We are hiring a Senior Python / AI Agent Engineer to work on that system with a clear path to owning it end to end: the agent graph, the retrieval layer underneath it, and the deterministic control layer around it.
That last part is the job. The hard problems here are not "can the model answer." They are: does the captured intent survive the handoff, does the tool get invoked on the right slot, does the state machine recover when the model drifts, and can you prove any of it with an evaluation suite rather than a vibe check. If you have shipped agents into production, you already know that the failures that hurt are the confident wrong answers, not the crashes.
You will work directly with our Tech Lead, Principal Engineer and Founder, and you will own decisions rather than tickets.
Read this section carefully, because it is the real job and it is what we will talk about if we speak.
A conversational agent and a product catalogue speak different languages. The agent can express far more than any storefront has pages for, and the translation between those two vocabularies is where agent products quietly break. Not with errors. With plausible answers that are wrong.
The layer that turns what a customer means into catalogue state, and keeps both vocabularies honest with each other, is the highest-value thing you will own here. Getting it right is mostly not a model problem. It is a contract problem, a state problem, and an evaluation problem.
If that sounds interesting rather than tedious, we should talk.
The agent layer
Retrieval
The resolution layer
Proving it works
You will also work in
Type: Full-time independent contractor, approximately 50 hrs/week. Overseas hires on a contractor agreement.
Term: 90-day initial contract with a strong path to extension based on results.
Compensation: $2,000 to $3,500 USD per month, weighted to demonstrated agent and retrieval depth. We are stating this openly so nobody wastes their time.
Location: Remote. Pakistan, the Philippines, and the wider South and Southeast Asia region.
Hours: Minimum 4 hours daily overlap with 9am to 5pm US Eastern Time.
Reports to: Tech Lead, with direct access to the Founder and Principal Engineer.
Start: Immediately.
Please answer these directly in your application. Substance matters more than length, and short specific answers beat long general ones. Applications without answers will not be reviewed.
An agent captures a user's intent correctly, the parameter is present and correct in the request, and the user still lands on the wrong result. Nothing errors and every test passes. Walk me through how you find that, and what check you add so it cannot happen again.
Describe an evaluation suite you built for an agent or a retrieval system. What did it assert, how often did it run, and what did it catch that a unit test would not have?
You have a multi-step conversational flow where the model sometimes skips a required slot. How do you make the flow reliable without making it feel scripted?
Tell me about a retrieval system you shipped on PostgreSQL and pgvector. What was the corpus, how did you chunk it, and how did you know your ranking was any good?
What are your current working hours, and what overlap can you commit to with 9am to 5pm US Eastern?
Put your answer to Question 1 at the very top of your message. Then send your resume or profile, a link to production work you can speak about in detail, and your answers to the remaining questions.
JustSoftLab Inc.