Software Engineer, Multimodal Backend Systems
eventualcomputing · San Francisco ·
- Employment
- Full time
- Category
- Backend
- Company size
- 11-50
eventualcomputing · San Francisco ·
ambient.ai · Redwood City
Baseten · San Francisco
Eventual · US
productboard · Prague, CZE
Builds core infrastructure for Eventual's multimodal AI data platform used by robotics/Physical AI teams — spanning real-time video streaming (WebRTC, GPU-accelerated codecs), LiDAR/sensor telemetry pipelines, or React/TypeScript data-curation frontends. Engineers work Rust, C++, Python, or Go on-site 4 days/week in San Francisco.
Robots and world models learn about the physical world from video: millions of hours of people cooking, building, carrying and fixing things. That footage is piling up faster than anyone can review it, and much of it is mislabeled, out of sync or not useful for training. Most teams still choose what to train on by having people watch a small sample. When a model learns from bad examples, it learns the wrong behavior.
Eventual is building data curation for Physical AI. We process video and sensor data at petabyte scale and extract signals from it: how hands and objects interact, what happens over time (a failed grasp, a retry, a recovery) and whether a clip is reliable enough to learn from. Leading Physical AI labs and robotics teams use these signals to decide what goes into their next training run.
The work spans large-scale data systems, computer vision and ML research, and it ships directly into how frontier models get trained. We’re a small team from AWS, Lyft and Tesla, backed by $30M from Felicis, CRV, Y Combinator and the co-founders of Databricks and Perplexity. We helped power the last generation of Physical AI in self-driving, and now we’re building what the next generation trains on.
Join our small (but powerful!) team, 4 days/week in our SF Mission District office.
Our goal is to build Scenario Mining and Data Curation for robot fleet data. We empower Physical AI and robotics teams to instantly find, curate, and stream the data they need to train frontier models.
Eventual is an agile team where every engineer has high ownership across the stack from our compute infrastructure, to our data storage/querying layers and model training/deployment.
As a Software Engineer working on our Multimodal Backend Systems, you will be responsible for building Eventual's core products and architecture. You will ship features that will be immediately used by our customers and will work with a tight-knit team that values open communication and cross-functional collaboration. We move quickly to solve a wide range of complex technical and product challenges. While we are an experienced team that can provide constant guidance and mentorship, we value engineers who can autonomously scope and solve difficult technical challenges.
We are seeking engineers with deep expertise in at least one of the following core domains:
1. Real-Time Video Infrastructure
Build our Streaming Architecture: Design, build, and optimize real-time WebRTC media pipelines and custom signaling mechanisms to stream multi-camera video feeds from robots to our platform
Manage Hardware-Accelerated Video Pipelines: Integrate and tune video codecs (H.264, HEVC/H.265, AV1) leveraging GPU acceleration (NVENC/NVDEC) for hardware decoding and dynamic bitrate adaptation
Scale Video Ingest & Storage: Engineer high-throughput video ingestion and distributed transcoding services that compress, index, and write petabyte-scale media corpora to object storage (AWS S3, GCS) for long-term retention and retrieval.
2. Robotics & Physical AI Telemetry
Build Spatial Data Pipelines: Construct specialized storage formats, spatial indexes, and processing workflows to ingest, align, and query dense 3D LiDAR point clouds, depth maps, and multi-camera spatial datasets.
Stream High-Frequency Telemetry at Scale: Develop high throughput ingestion pipelines capable of capturing, parsing, and storing real-time sensor streams and high-frequency robot state telemetry across our customers’ fleets.
Ensure Sensor-to-Video Synchronization: Coordinate time-sync protocols (PTP/NTP) across disparate sensor feeds to temporally align LiDAR point clouds, IMU telemetry, and video frames into unified data structures for downstream consumption.
3. Product & Multimedia Frontend
Develop Data-Dense User Interfaces: Architect intuitive, high-performance web applications using React, TypeScript, and modern state management to handle continuous streams of data.
Visualize 3D Spatial & Video Streams: Build custom frontends to render 3D LiDAR point clouds, spatial bounding boxes, and multi-camera video feeds.
Build Agentic/AI-Native Product Workflows: Create agentic data-exploration, curation and search tools that empower robotics operators and AI researchers to review, annotate, and analyze complex physical AI datasets.
We are looking for strong engineers who are problem-solvers at heart—combining excellent coding and architectural fundamentals in languages like Rust, C++, Python, or Go with a drive to reach for lower-level primitives when performance and efficiency demand it.
In-person tight knit team with 4x a week in office
Competitive comp and startup equity
Catered lunches and dinners for SF employees
Commuter benefit
Team building events & poker nights
Health, vision, and dental coverage
Flexible PTO
Latest Apple equipment
401k plan with match!
Writer · San Francisco, CA