Senior Site Reliability Engineer
- Seniority
- Senior
- Work model
- Remote · LatAm-friendly
- Location
- Remote (Brazil)
- Posted
- 20d ago
<p><strong>About SecurityScorecard:</strong></p> <p><span style="font-weight: 400;">SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries. Founded in 2013 by security and risk experts Dr. Alex Yampolskiy and Sam Kassoumeh and funded by world-class investors, SecurityScorecard’s patented rating technology is used by over 25,000 organizations for self-monitoring, third-party risk management, board reporting, and cyber insurance underwriting; making all organizations more resilient by allowing them to easily find and fix cybersecurity risks across their digital footprint. </span></p> <p><span style="font-weight: 400;">Headquartered in New York City, our culture has been recognized by Inc Magazine as a "Best Workplace,” by Crain’s NY as a "Best Places to Work in NYC," and as one of the 10 hottest SaaS startups in New York for two years in a row. Most recently, SecurityScorecard was named to Fast Company’s annual list of the</span><a href="https://securityscorecard.com/blog/blog-top-innovators-of-the-year/"><span style="font-weight: 400;"> </span><span style="font-weight: 400;">World’s Most Innovative Companies for 2023</span></a><span style="font-weight: 400;"> and to the Achievers 50 Most Engaged Workplaces in 2023 award recognizing “forward-thinking employers for their unwavering commitment to employee engagement.” SecurityScorecard is proud to be funded by world-class investors including Silver Lake Waterman, Moody’s, Sequoia Capital, GV and Riverwood Capital.</span></p> <p><strong><br>About the Team:</strong></p> <p>As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization of our Kubernetes-based infrastructure and CI/CD systems. You will also own the infrastructure behind our AI tooling — building MCP servers and defining safe, auditable AI access patterns for production systems. You'll work hands-on with engineering teams to accelerate delivery, ensure production reliability, and embed best practices for automation, observability, and resilience.</p> <p><strong><br>About the Role:</strong></p> <ul> <li>Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.</li> <li>Build and operate AI tooling infrastructure — stand up MCP servers and establish secure, governed AI access and guardrails for production systems.</li> <li>Optimize and maintain CI/CD pipelines, improving reliability, speed, and rollback safety.</li> <li>Implement progressive delivery strategies such as blue/green and canary deployments.</li> <li>Advance Infrastructure as Code with Terraform, Helm, and Argo CD, defining reusable patterns for the org.</li> <li>Operate and optimize streaming and analytics infrastructure: Kafka, Flink, and ClickHouse.</li> <li>Build automated testing into the CI/CD lifecycle.</li> <li>Improve system observability — define SLOs, alerts, and dashboards.</li> <li>Lead incident response and postmortems, focusing on root cause and durable fixes.</li> <li>Mentor engineers across teams on Kubernetes, CI/CD, and cloud infrastructure.</li> </ul> <p><strong><br>Required Qualifications:</strong></p> <ul> <li><strong>6+ years</strong> in <strong>SRE</strong>, <strong>DevOps</strong>, or <strong>Infrastructure roles</strong>, with significant <strong>production Kubernetes experience</strong>.</li> <li>Hands-on experience integrating <strong>AI/LLM tooling</strong> into engineering or operational workflows (e.g., <strong>MCP servers</strong>, AI agents acting on infrastructure), and a clear grasp of the <strong>security and governance</strong> considerations of giving AI access to production.</li> <li>Proven success building CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, or similar).</li> <li>Strong with <strong>Kubernetes internals</strong> and <strong>managed services</strong> like <strong>EKS</strong>, <strong>GKE</strong>, or <strong>AKS</strong>.</li> <li>Expertise with <strong>Infrastructure as Code</strong> (<strong>Terraform</strong>, <strong>Helm</strong>, <strong>Pulumi</strong>) and <strong>GitOps</strong>.</li> <li>Proficient in <strong>Python</strong>, <strong>Bash</strong>, or <strong>Go</strong>.</li> <li>Knowledge of <strong>observability tooling</strong> (<strong>Prometheus</strong>, <strong>Grafana</strong>, <strong>Datadog</strong>, <strong>OpenTelemetry</strong>).</li> <li>Production experience with <strong>Kafka</strong>, <strong>Flink</strong>, and <strong>ClickHouse</strong>.</li> <li><strong>Strong communication</strong> and <strong>cross-team collaboration</strong> skills.</li> </ul> <p><strong><br>Preferred Qualifications:</strong></p> <ul> <li><strong>Multi-region</strong> or <strong>multi-cluster</strong> Kubernetes experience.</li> <li><strong>Chaos engineering</strong> or <strong>resilience testing</strong>.</li> <li><strong>Security scanning</strong>, <strong>compliance automation</strong>, or <strong>policy-as-code</strong>.</li> <li><strong>LLM observability/tracing tooling</strong> (<strong>Langsmith</strong>, <strong>Langfuse</strong>) or <strong>MLOps workflows</strong>.</li> <li>Contributions to <strong>open-source</strong> Kubernetes or CI/CD projects.</li> </ul> <p><span data-sheets-value="{"1":2,"2":"SecurityScorecard is an industry-leading cybersecurity company backed by Google, Sequoia, and Riverwood. Our mission is to make the world a safer place. We measure your and your vendors' cyber-health by assigning a security rating of \"A\" through \"F\" based on outside-in, non-intrusive data. Our Comprehensive security ratings, advanced data analytics, and actionable insights discover Third-Party Vulnerabilities & Security Gaps In Real-Time. \nHeadquartered in NYC with over 200+ employees globally, raised over $110M USD, used by 1,000+ enterprise customers, and rating 1.5 million companies. We have created a new category of enterprise software, and our cul