About the Role
The NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our mission is to build highly scalable and efficient infrastructure for open-source AI at a global scale. We are seeking an expert in high-performance LLM serving systems and inference optimization to push the boundaries of how large language models are served.
Responsibilities
- Architect and maintain production high-traffic LLM serving systems.
- Optimize throughput, latency, and cost for leading open-source LLMs.
Requirements
- Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT.
- Deep knowledge of state-of-the-art GPU architectures, and effectively exploit them using PyTorch, Triton, CuTe, CUDA, etc.
- Proven track record in designing and maintaining end-to-end high-traffic LLM serving systems.
- Strong problem-solving skills and ability to communicate technical ideas clearly.
Skills
- LLM inference
- SGLang
- vLLM
- TensorRT
- GPU architectures
- PyTorch
- Triton
- CuTe
- CUDA
- LLM serving systems
- Trusted Execution Environments (TEE)
- Open-source LLM inference engines
Location
- San Francisco
- Remote
Work Type
- Remote
About the Company
- NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI.
- Mission is to build highly scalable and efficient infrastructure for open-source AI at a global scale.
