Found Description
Responsibilities
- Develop and optimize LLM serving systems based on DeepX NPU
- Design and implement runtime and inference engines for LLM
- Analyze LLM serving performance and resolve bottlenecks
Frontier of On-device AI Semiconductors
for Everyone, Everywhere
Qualifications
- Bachelor’s degree or higher in Computer Science or a related field
- Experience in software development using C/C++ and Python
- Understanding of LLMs or deep learning frameworks (e.g., TensorFlow, PyTorch)
- Experience with Linux-based development environments
- Basic knowledge of computer architecture and parallel processing
Preferred Qualifications
- Experience with AI accelerator hardware such as NPUs or GPUs
- Hands‑on experience with model compilers/runtimes like ONNX, TVM, or TensorRT
- Knowledge of model optimization techniques such as quantization and pruning