Found Description
Company
We’re working with a fast-growing startup building a next‑generation AI compiler focused on accelerating deep learning training through hardware‑aware optimisation and LLM‑driven code generation. You’ll join a small, highly technical team working at the intersection of compilers, ML systems, and performance engineering, with real ownership over core compiler components.
Responsibilities
- Build and optimise AI compiler components for training workloads
- Develop compiler passes, runtimes, and hardware backends
- Apply framework‑ and kernel‑level performance optimisations
- Work closely with ML, compiler, and LLM‑focused engineers
Qualifications
- Strong Python and C/C++ skills
- Experience with CUDA, LLVM, Triton, or custom codegen
- Familiarity with PyTorch, TensorFlow, ONNX, or TensorRT
- Experience tuning performance on GPUs or accelerators