Found Description
We are looking for a highly skilled AI / LLM Engineer to lead the training, alignment, and optimization of large language models. This role focuses on Reinforcement Learning from Human Feedback (RLHF) and end-to-end post-training pipelines, while ensuring models are efficient and production-ready.
You will play a key role in bridging AI alignment research and systems engineering , driving innovation in model performance, safety, and scalability.
Key Responsibilities
- Lead and manage the end-to-end RLHF pipeline (data collection, reward modeling, RL fine-tuning – PPO, DPO, GRPO, RLAIF)
- Design and implement Supervised Fine-Tuning (SFT) pipelines using models like LLaMA, Mistral, and Qwen
- Build and train reward models based on human feedback
- Develop annotation pipelines (guidelines, calibration, dataset curation)
- Apply Constitutional AI & RLAIF to red...