At Capgemini Engineering, the world leader in engineering services, we bring together a global team of engineers, scientists, and architects to help the world's most innovative companies unleash their potential. From autonomous cars to life-saving robots, our digital and software technology experts think outside the box as they provide unique R&D and engineering services across all industries.
Requirements
- Master's or PhD degree in Data Science, Computer Science, Applied Mathematics, Statistics, Physics, or a related field
- 2+ years of hands-on experience working with machine learning datasets
- Experience with time-series, sensor, image, or video data
- Strong Python skills and experience with NumPy, pandas, matplotlib, and seaborn
- Experience with dimensionality reduction and representation analysis techniques such as UMAP, t-SNE, and PCA
- Solid understanding of machine learning fundamentals, model evaluation, and diagnostics
- Experience supporting both traditional machine learning and deep learning projects
Nice to Have
- Experience working with sensor data, including radar, magnetic, environmental, 3D, or IoT datasets
- Familiarity with scikit-learn preprocessing workflows
- Experience handling imbalanced datasets, noisy labels, sensor noise, and data drift
- Knowledge of model interpretability, feature importance, and embedding analysis
- Experience working with data annotation and labeling teams
- Familiarity with MLflow, Weights & Biases, DVC, or similar tools
Responsibilities
Data Understanding and Representation Analysis
- Analyze high-dimensional sensor and feature datasets using UMAP, t-SNE, PCA, and similar techniques
- Identify clusters, anomalies, blind spots, distribution gaps, and class or environment mismatches
- Diagnose dataset shift, domain drift, sparsity, and representation collapse
Model-Aware Data Analysis
- Perform data analysis aligned with classical ML models including XGBoost, SVR, k-NN, and tree-based models
- Support analysis for deep learning models such as CNNs and Transformers
- Analyze embeddings, confusion matrices, and model failure patterns to trace errors back to data issues
Data Quality and Curation
- Investigate imbalanced data, noisy sensor signals, mislabeled samples, and ambiguous cases
- Develop approaches for improving weakly labeled or unlabeled data including clustering and pseudo-labeling
- Perform data mining on large collections of field data to extract insights and patterns
- Design processes for converting noisy or partially verified data into high-quality validated datasets
Insight-Driven Improvements
- Translate exploratory findings into clear recommendations for data filtering, relabeling, or new data collection
- Advocate for and implement data-centric improvements to enhance model robustness
- Work closely with engineering teams to integrate improved data workflows into ML pipelines
What You Will Love About Working Here
- Health insurance from the first days, regardless of the probationary period
- The Gift from the Company - Christmas holidays from December 25 to December 31
- Cooperation with Superhumans center and Veteran HUB with support for psychological counseling and company policies friendly to military and veterans