Skip to content

Repository files navigation

Awesome SRE Awesome

You want your computer systems to run well, and the subjective definition of what well means depends on the nature of the system and your goals regarding it.

Most of the time, the primary motivation for companies is to create profit for the owners and shareholders.

The definition of running well will therefore be a derivative of the business model objectives.

"Hope is not a strategy."

Contents

1. Site Reliability Engineering

2. SRE Culture

3. DevOps

4. Monitoring and Observability

5. Alerting

6. Incident Response and Post-Mortem

  • A collection of post-mortems - Curated collection of post-mortems from various companies and incidents.
  • A collection of postmortem templates - Collection of templates for writing effective post-mortems.
  • Our incident postmortem template - Hosted Graphite postmortem template.
  • Postmortem exercise - Google's postmortem exercise document.
  • incident.io - Incident management platform that helps teams respond, communicate, and learn from incidents directly within Slack.
  • Squadcast - SRE platform for incident management, on-call scheduling, and reliability workflows.
  • PagerDuty - Digital operations management platform for real-time incident response.
  • Splunk On-Call - On-call and incident response tool by Splunk, formerly known as VictorOps.
  • OpsGenie - On-call and alert management to keep services always on.
  • AlertOps - Transforms real-time operational intelligence into automated incident response.
  • Blameless - SRE platform for incident management, retrospectives, and reliability insights.
  • OnPage - Incident alert management system with a secure smartphone app for response teams.
  • PagerTree - Intelligent alert routing for the modern team.
  • Cabot - Get alerted when services go down or metrics go crazy.
  • xMatters - Service reliability platform to automate operations workflows.
  • Derdack Enterprise Alert - Enterprise alert notification software.
  • Bigpanda - AIOps event correlation and automation platform.
  • OpenDuty - (Deprecated) Open source incident escalation tool similar to PagerDuty, no longer maintained.
  • ngDesk - All-in-one application that includes support, sales, asset management, marketing and pager.
  • Geneos - Real-time monitoring for all your environments in one platform.
  • FireHydrant - Tools for service catalogs, incident response, status pages, and retrospectives.
  • Rootly - Incident management platform with automated workflows and Slack integration.

7. On-Call

8. Chaos Engineering

9. Automation and Toil Reduction

  • Eliminating Toil - Google SRE Book - Chapter on identifying and reducing toil in SRE practice.
  • Rundeck - Open source runbook automation for incident management, business continuity, and self-service operations.
  • Ansible - Simple, agentless IT automation platform for configuration management, application deployment, and orchestration.
  • Terraform - Infrastructure as Code tool for building, changing, and versioning infrastructure safely and efficiently.
  • Pulumi - Infrastructure as Code using familiar programming languages like Python, Go, JavaScript, TypeScript, and C#.
  • Shoreline - Incident automation platform that enables on-call engineers to debug and repair production issues with real-time automation.
  • StackStorm - Open source event-driven platform for runbook automation, ChatOps, and auto-remediation.

10. Capacity Planning

11. Runbooks and Playbooks

12. Performance

13. SLOs and SLIs Tools

  • SLO Generator - Tool by Google to compute and export Service Level Objectives, Error Budgets and Burn Rates using YAML configurations.
  • SLO Computer - Simplifies computing SLOs, error windows and alerts.
  • SLO Tracker - A simple but effective way to track SLOs and Error budgets with webhook integration for alerting tools.
  • SLO exporter - Computes standardized SLI and SLO metrics based on events from various data sources.
  • Pyrra - Makes SLOs with Prometheus manageable, accessible, and easy to use for everyone.
  • Sloth - Easy and simple Prometheus SLO generator.
  • Nobl9 - SLO platform that connects to any data source to provide reliability insights and error budget tracking at enterprise scale.
  • OpenSLO - Open specification for defining and interfacing with SLOs, allowing for a common, vendor-agnostic approach to reliability tracking.
  • NthLayer - Reliability requirements as code, generating Grafana dashboards, Prometheus alerts, SLOs, and PagerDuty configs from service.yaml.

14. Books

15. Examples and Sandboxes

  • Observability Sandbox - Get up and running with Prometheus, Thanos, Grafana, and more using Docker and Docker Compose.
  • School of SRE - LinkedIn's comprehensive self-study curriculum covering fundamentals to advanced SRE topics.
  • SRE University - Curated list of courses and resources for learning SRE.
  • Google SRE Resources - Official Google resources including talks, blog posts, and case studies.
  • Operate First - Community-driven SRE practices for open source cloud operations.

16. Community and Forums

  • SREcon - USENIX conference dedicated to Site Reliability Engineering.
  • CNCF TAG Observability - CNCF Technical Advisory Group for observability topics.
  • Google SRE Resources - Official talks, blog posts, and educational content from Google SRE teams.
  • SRE Weekly - Weekly newsletter curating the best SRE news and articles.
  • Awesome SRE - A curated list of awesome Site Reliability and Production Engineering resources.

17. References

18. License

CC0

19. Contributing

Contributions welcome! Read the contribution guidelines first.

Thank you!

About

Awesome SRE page

Topics

Resources

Code of conduct

Contributing

Stars

14 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors