Core Responsibilities

Lead the technical strategy for open source model inference, defining serving architecture and cost optimization strategies. Design and implement solutions for token compression, prompt caching, and KV-cache optimization to enhance throughput and reduce costs. Develop core components of the gateway in Go, ensuring high availability and horizontal scalability. Mentor team engineers and review technical designs to elevate code quality standards.

Requirements

Experience leading technical initiatives or software engineering teams. Hands-on experience with open source LLM model inference. Proficiency in Go for high-performance system development, including concurrency, gRPC, and memory management. Knowledge of GPU architecture, CUDA, and hardware considerations for inference processes.

Additional Information

Experience Level

Lead / Principal

Job Language

Spanish

Employment Type

Full-time

Work Mode

Hybrid