
// Open Role at Chess.com
Join Chess.com as a Senior Site Reliability Engineer to ensure the stability and performance of a global gaming platform serving over 250M users. Lead infrastructure design, cloud migration, and automation initiatives for a mission-driven, fully remote team.
About Us
Chess.com is one of the largest gaming sites in the world and the #1 platform for playing, learning, and enjoying chess.
We are a team of 600+ fully remote people in 60+ countries working hard to serve the global chess community. We are here to support 250M+ chess players worldwide with the best possible product, content, and tools to serve the community!
We are a tech company. A gaming company. A content company. And we do it all with passion and commitment to the game. Above all we prize our mission-driven, flat, life-celebrating, no-corporate culture, and we look forward to meeting you and learning more about what you can bring to the team.
About the role
The Site Reliability Engineer will play a critical role in ensuring the stability, performance, and scalability of our global gaming platform infrastructure. This position exists to bridge the gap between development and operations, maintaining high availability for millions of concurrent users while supporting rapid feature development and deployment. The SRE will be instrumental in building resilient systems that can handle massive scale across multiple regions, directly impacting user experience and platform reliability.
As our platform continues to grow and serve a global community, this role will drive the technical infrastructure decisions that enable seamless gaming experiences. The position requires both deep technical expertise and collaborative leadership to work across engineering teams, ensuring our systems can scale efficiently while maintaining the performance standards our users expect.
What you'll do
Preferred Skills
About the Opportunity
---
You can learn more about us here:
Design and implement multi-regional resilient infrastructure for millions of concurrent sessions. Lead hybrid cloud migration strategy. Own on-call rotation and incident response. Architect monitoring and alerting systems. Collaborate on infrastructure-as-code and deployment pipelines. Optimize system performance through capacity planning and load testing. Establish and maintain security protocols. Partner with engineering teams on scalable solutions. Drive automation initiatives. Mentor team members on SRE best practices. Bachelor's degree in Computer Science or equivalent practical experience. 5+ years in SRE, DevOps, or infrastructure engineering. Experience managing bare-metal server infrastructure. Strong proficiency with UNIX/Linux. Experience with cloud platforms and infrastructure-as-code tools. Hands-on experience with configuration management systems. Solid understanding of networking fundamentals. Experience with containerization and orchestration technologies. Proficiency with monitoring and observability tools. Experience with relational and NoSQL databases, including performance optimization. Strong collaboration and communication skills. Demonstrated sense of ownership and accountability.