Found Description
About The Company
An early-stage AI infrastructure startup specializing in inference optimization. The company develops software to increase compute performance per watt on GPUs, integrating with major serving stacks to provide kernel-level power telemetry. By utilizing a proprietary optimization engine and runtime controller, the platform helps AI teams maintain high throughput under power constraints, dynamically adjusting performance as models and hardware conditions evolve.
About The Role
As a Senior Performance Engineer and founding team member, you will architect solutions across the entire inference stack. You will have direct ownership over optimizing kernels, shaping the serving engine, and improving orchestration to redefine large‑scale model deployment. Your work will directly impact the cost and energy efficiency of AI at scale, pushing GPU utilization toward theoretical limits. This role offers the opportunity to translate comp...
Ready to Apply?
Submit your application for Senior Performance Engineer - LLM Inference Optimization at Amplify.LA
Apply Now