Constellation is hiring a senior software engineer to build the backend behind its data collection, owning the hybrid on-premise and cloud architecture handling TB-scale multimodal ingestion daily.
Constellation is building a foundation model of human state, translating multimodal human data into generalizable embeddings to power clinical, robotic and AI systems. It is a venture-backed team of 15 to 20 researchers and engineers based in San Francisco's Mission District.
About the role
This role builds the backend for Constellation's data collection, owning the architecture that captures, buffers, processes, stores and retrieves its multimodal data. You will design a hybrid on-premise and cloud infrastructure supporting TB-scale data collected daily, working with the AI team to deliver high-performance, training-ready datasets.
What you will do
- Architect a fault-tolerant data ingestion engine handling intermittent sensor connectivity, hardware failures and network issues without data loss.
- Write robust multi-process applications handling concurrent data streams.
- Architect a unified device abstraction layer ingesting data from heterogeneous peripherals and network interfaces including USB, BLE and TCP/IP.
- Optimize I/O bound operations and track bandwidth and latency across the infrastructure.
- Define the storage topology for TB-scale daily ingestion, selecting file formats for video, audio, text and time-series and designing schemas optimized for rapid indexing and retrieval.
- Manage provisioning and configuration of on-premise servers, and build the synchronization logic moving terabytes from local edge buffers to the central cloud repository.
What they are looking for
- 3+ years shipping production-grade backend systems, with expert proficiency in Python, Rust, C++ or Go.
- Proficiency with high-performance messaging and queuing tools such as ZeroMQ, Kafka, RabbitMQ or Redis Pub/Sub.
- Strong experience architecting infrastructure on AWS, GCP or Azure.
- Low-latency streaming protocols such as WebRTC, RTSP or HLS, and programmatic video processing with FFmpeg, GStreamer or OpenCV.
- High-performance data transfer techniques including zero-copy networking, shared memory and memory-mapped files.
- Time-series databases such as TimescaleDB, InfluxDB or ClickHouse, and data lake formats including Parquet, Avro and Arrow.