Designs, builds, and owns secure sandboxed infrastructure for sensitive AI model evaluations at Reflection, including controlled data pipelines, access controls, audit logging, and evaluation orchestration tooling, with production-quality Python as the core technology.
You will design, build, and own secure infrastructure for sensitive model evaluations. You will create sandboxed environments, controlled data pipelines, access controls, audit logging, evaluation orchestration tooling, and capability-measurement systems. You will improve platform reliability, security, and developer experience.
Responsibilities
Design and build secure sandboxed infrastructure for sensitive model evaluations
Build controlled data pipelines and storage for sensitive evaluation material
Translate evaluation designs into reliable and scalable systems
Build evaluation orchestration tooling and harnesses
Develop infrastructure to measure AI capability uplift
Implement guardrails monitoring and compartmentalization
Write production-quality Python for evaluation systems
Improve platform reliability security and developer experience
Requirements
Software engineering expertise in Python
Experience building secure sandboxed or isolated execution environments
Knowledge of least privilege need-to-know access RBAC secrets management encryption audit logging and defense in depth
Experience building data pipelines for sensitive or restricted data
Ability to own ambiguous cross-functional problems end to end
Discretion integrity and sound judgment for sensitive projects
Benefits
Stock options
Medical dental vision and life insurance
Annual wellness allowance
Daily office lunch and dinner
22 weeks of paid parental leave
Unlimited paid time off in the United States
30 days of vacation in the United Kingdom
Visa sponsorship support
Regular off-sites happy hours and team celebrations