Core Responsibilities

Lead structural initiatives to improve resilience, uptime, and scalability. Design secure and resilient architectures tailored to business needs, integrating AI-based solutions for system reliability. Model and prioritize uptime based on service criticality, and evaluate technical maturity to design evolution plans. Drive advanced observability practices and identify technical bottlenecks through deep profiling and troubleshooting.

Requirements

5+ years of experience in development, with expertise in distributed systems, microservices, scalable APIs, and cloud architectures. Deep knowledge of uptime management, incident response, and SRE engineering practices. Advanced proficiency in observability and troubleshooting tools such as Datadog, New Relic, and Kibana, along with experience in AI/ML solutions applied to operations.

Additional Information

Experience Level

Senior

Job Language

Spanish

Employment Type

Full-time

Work Mode

Hybrid