
// Open Role at Penn Interactive
Lead critical incident management for Penn Entertainment's cutting-edge online gaming and sports media platforms, driving resolution and service improvements across North America with a focus on SRE and devops excellence.
PENN Entertainment, Inc. is North America’s leading provider of integrated entertainment, sports content, and casino gaming experiences. From casinos and racetracks to online gaming, sports betting and entertainment content, we deliver the experiences people want, how and where they want them.
We’re always on the lookout for those who are passionate about creating and delivering cutting-edge online gaming and sports media products. Whether it’s through Hollywood Casino, theScore Bet Sportsbook, or theScore media app, we’re excited to push the boundaries of what’s possible. These state-of-the-art platforms are powered by proprietary in-house technology, a key component of PENN’s omnichannel gaming and entertainment strategy.
When you join PENN Entertainment’s digital team, you’ll not only work on these cutting-edge platforms through theScore and PENN Interactive, but you’ll also be part of a company that truly cares about your career growth. We’re committed to supporting you as you expand your skills and explore new opportunities.
With locations throughout North America, you can build a future at PENN Entertainment wherever you are. If you want to challenge conventions in gaming, media and entertainment, we want to talk to you.
**About the Role & Team
**As part of the team, you will be working with a team of smart, friendly, and dedicated Engineers, Product Managers and Designers determined to deliver some of the best apps the market has to offer. We are looking for an Incident Commander to join our site reliability team, to work cross-functionally across engineering, and be the front line for incidents and working with Release Engineering to help prevent new events.
This is a position responsible for all incidents across various organizations within the company, both online and physical, which includes P1, P2, P3 and P4. Classifying and documenting all incidents and carrying out support, assisting and driving all incidents, regarding investigation, hierarchical and technical escalation, diagnosis and recovery and root cause analysis. Additionally driving improvements to our service delivery and release processes based on disruption reports.
About the Work
• Drive and enhance collaboration with other Incident Commanders, Customer Support, Application and Engineering teams - cross-functional teams to lead real-time incident management.
• Provides Leadership for developing Practices, Frameworks, Process Flows, Templates and Process Guides
• Continuously improve and enhance the internal framework, methodology, processes, and tools
• Developing and maintaining key practical capabilities
• Collaborating with SRE Teams and Infrastructure teams to identify requirements and gaps resulting in downtime or blindspots.
• Recommends innovative solutions that enable the organization to deliver on its objectives and goals.
• Promote opportunities for Continuous Service Improvements
• Manage and update Root Cause Analysis documentation.
• Lead SRE communications to stakeholders via E-mail, Slack, & Teams in timely manner
• Lead initiatives to promote JIRA Release Ticket management, quality and alignment with Incident management communication supporting SLAs
• Other duties as required.
About You
• Experience in a similar role or incident management role.
• Experience and understanding of Containerization (Docker & Kubernetes preferred)
• Automation: Understanding of configuration management and infrastructure as code tools. Terraform, Ansible, Helm, etc.
• Experience with a programming language.
• Comfortable within Linux environments and needs.
• Experience working with AWS, GCP, and on-premises environments.
• Ability to work independently and learn quickly with little supervision.
• Ability to handle multiple projects simultaneously.
• Willingness to drop everything and take on an ad-hoc task.
• You’re the type of individual who is tech-savvy and passionate about learning new technologies and tools.
• Out-going, and able to keep a conversation going naturally to extract needed information
• A degree in computer science, engineering, and/or similar experience.
• Nice to have: Postgres, MySQL, Elastic Search, Kafka, Redis, Terragrunt, Prometheus, Python, Talos Linux
What We Offer
Competitive compensation package.
Fun, relaxed work environment.
Education and conference reimbursements.
Opportunities for career progression and mentoring others.
#LI-REMOTE
Salary Range
$90,000—$135,000 USD
_Penn Interactive is proud to be an equal opportunity workplace. We will consider all qualified applicants for employment without regard to race, color, religion, age, sex, sexual orientation, gender identity, national origin, disability, veteran status, genetic information, or any other basis protected by applicable law.Base pay is one part of the Total Rewards that Penn Interactive provides to compensate and recognize employees for their work. Most sales positions are eligible for a Commission under the terms of an applicable plan, while most non-sales positions are eligible for a Bonus. Additionally, Penn Interactive provides best-in-class benefits to eligible employees. We believe that benefits should connect you to the support you need when it matters most, and should help you care for those who matter most. That’s why we provide an array of options, expert guidance and always-on tools, that are personalized to meet the needs of your reality – to help support you physically, financially and emotionally through the big milestones and in your everyday life.
_
- Experience in a similar role or incident management role. - Experience and understanding of Containerization (Docker & Kubernetes preferred). - Understanding of configuration management and infrastructure as code tools: Terraform, Ansible, Helm, etc. - Experience with a programming language. - Comfortable within Linux environments and needs. - Experience working with AWS, GCP, and on-premises environments. - Ability to work independently and learn quickly with little supervision. - Ability to handle multiple projects simultaneously. - Willingness to drop everything and take on an ad-hoc task. - Tech-savvy and passionate about learning new technologies and tools. - Out-going, and able to keep a conversation going naturally to extract needed information. - A degree in computer science, engineering, and/or similar experience. - Drive and enhance collaboration with other Incident Commanders, Customer Support, Application and Engineering teams - cross-functional teams to lead real-time incident management. - Provide Leadership for developing Practices, Frameworks, Process Flows, Templates and Process Guides. - Continuously improve and enhance the internal framework, methodology, processes, and tools. - Develop and maintain key practical capabilities. - Collaborate with SRE Teams and Infrastructure teams to identify requirements and gaps resulting in downtime or blindspots. - Recommend innovative solutions that enable the organization to deliver on its objectives and goals. - Promote opportunities for Continuous Service Improvements. - Manage and update Root Cause Analysis documentation. - Lead SRE communications to stakeholders via E-mail, Slack, & Teams in timely manner. - Lead initiatives to promote JIRA Release Ticket management, quality and alignment with Incident management communication supporting SLAs.