We're looking for a Data Scientist to build and ship ML models on top of our data platform, working closely with Data Engineers and Software Engineers. What You'll Do:
Design, train, and evaluate ML models
Define feature requirements and partner with Data Engineering to source them from our pipelines
Own model deployment, serving, and production monitoring (drift, degradation, retraining)
Explore large datasets directly via Spark/SQL rather than relying only on pre-built extracts
Communicate findings and model behavior to technical and non-technical stakeholders
What You'll Need:
Strong Python: pandas, numpy; ML and LLM models; RAG systems and AI generative approaches; experience with PyTorch or TensorFlow
Hands-on PySpark for feature engineering at scale
Solid SQL skills against large tables
Experience owning the full model lifecycle: training → deployment → monitoring
Understanding of production model serving (batch, REST, or streaming)