SRE Made Simple : Master reliability through observability and automated infrastructure as code

Description: Site reliability engineering is the modern approach to improving the reliability of software systems. As systems grow with more features and users, issues and outages become more common, often leading to revenue loss. This book explores SRE practices, along with the design patterns and...

Description complète

Enregistré dans:
Détails bibliographiques
Auteur principal: Kumar, Jayant
Format: Livre numérique
Langue:Anglais
Publié: New Delhi : BPB Publications 2026.
Paris : Cyberlibris
Accès en ligne:Accès Université d'Orléans et IFPM
Note: Couverture. https://static2.cyberlibris.com/books_upload/300pix/9789378549076.jpg
Cyberlibris (ScholarVox) corpus Informatique
Autres localisations: Voir dans le Sudoc
Edition sous un autre format:• SRE Made Simple, Master reliability through observability and automated infrastructure as code, Jayant Kumar, New Delhi, BPB Publications, 2026, 1 vol. (417 p.), 978-93-7854-907-6
Description
Résumé:Description: Site reliability engineering is the modern approach to improving the reliability of software systems. As systems grow with more features and users, issues and outages become more common, often leading to revenue loss. This book explores SRE practices, along with the design patterns and tools that can be used to enhance system reliability. In this book, the mindset of an SRE engineer will be explored, and the evolution of team culture required to support SRE will be discussed. Readers will understand the metrics that need to be tracked for SRE, along with the sub-practices adopted to improve site reliability. The building blocks of site reliability engineering will be outlined. Readers will also explore the actions involved in implementing SRE across software engineering. Some tools used to implement SRE practices will also be introduced. Additionally, real-world examples will be included to provide practical understanding. This book will prepare readers towards the implementation and adoption of SRE practices within their team and organization. It will also help them understand their existing SRE practices and guide them to improve them further. For readers new to the concept of SRE, this book will help them understand what SRE is and how it should be implemented. What you will learn: Manage SRE error budget metrics and scale across organizations; Define SLI, SLO, and SLA metrics and manage SRE error budgets effectively; Optimize latency and system throughput; Utilize AIOps for predictive incident detection; Understanding incident management and modern release engineering practices; Explore tools and understand how AI helps SRE in improving site reliability. Who this book is for: This book is for DevOps engineers, software architects, and technical managers seeking to master reliability. While beneficial for senior executives, readers should possess a foundational understanding of software lifecycles and infrastructure to successfully adopt SRE practices that optimize business revenue
Description:Couverture. https://static2.cyberlibris.com/books_upload/300pix/9789378549076.jpg
Cyberlibris (ScholarVox) corpus Informatique
Accès:L'accès en ligne est réservé aux établissements ou bibliothèques ayant souscrit l'abonnement. Cyberlibris