Atos is hiring a Lead AI Engineer to build and operate the model, infrastructure, and data layer of the Belgium AI delivery practice. The role trains and fine-tunes custom large and small language models, builds and administers AI infrastructure, and leads data management, modelling, and integration work for AI projects.
Key Responsibilities
- Build and train custom LLMs and SLMs.
- Fine-tune, align, and adapt models for project needs.
- Create and administer AI infrastructure, including compute, serving, orchestration, autoscaling, and cost controls.
- Lead data management for AI projects, including pipelines, vector stores, feature stores, and lineage tracking.
- Lead modelling decisions across the portfolio, selecting foundation models, fine-tuning, or custom SLM approaches.
- Lead AI integration into project codebases through APIs, retrieval layers, and agent frameworks.
- Run evaluation harnesses for trained models, defining metrics, test sets, and CI-based quality gates.
- Partner with the AI Governance and Platform Lead on governance documentation.
- Mentor delivery-team engineers on model selection, fine-tuning, and data-quality practices.
Required Qualifications
- Bachelor’s or Master’s degree in Computer Science, Machine Learning, Engineering, or equivalent experience.
- 6+ years in ML engineering, deep learning, or AI platform engineering.
- 3+ years training, fine-tuning, or deploying LLMs in production.
- Experience with LoRA, QLoRA, instruction tuning, RLHF, or DPO.
- Knowledge of model-serving runtimes and GPU infrastructure on a major cloud platform.
- Strong data engineering experience with pipelines, vector stores, feature stores, and lineage tracking.
- Senior-level Python skills with PyTorch and modern AI tooling.
- Understanding of evaluation methodology, hallucination measurement, and retrieval failure analysis.
- Production experience taking at least one significant model from training to live operation.
Preferred Qualifications
- Experience training small language models or running continued pretraining.
- Hands-on RAG experience at scale.
- Experience with sovereign EU AI stacks.
- Hands-on experience with agent frameworks in production.
- AI FinOps experience covering cost control, right-sizing, caching, and workload optimisation.
- Open-source contributions, papers, or talks.
Technical Skills
PyTorch, Hugging Face, DeepSpeed, Accelerate, vLLM, TGI, LoRA/QLoRA, DPO, RLHF, Triton, Ray Serve, SageMaker, AWS Bedrock, Databricks, Snowflake, vector stores, feature stores, knowledge graphs, AWS, Azure, GCP, Kubernetes, Docker, GPU infrastructure, Langfuse, Arize Phoenix, Ragas, DeepEval, Promptfoo, OpenAI Evals, Python, TypeScript, Go, Rust, CI/CD, and infrastructure-as-code.