『The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering』のカバーアート

The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering

The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering

著者: Fexingo
無料で聴く

Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your system is already in production? Lucas and Luna don't pretend to have final answers — they build the conversation so you can draw your own. If you've ever argued about whether a page was necessary or whether an SLO should be tightened, this is your show. #SiteReliabilityEngineering #SRE #Uptime #ProductionEngineering #IncidentResponse #ErrorBudgets #SLOs #Postmortem #ToilAutomation #CapacityPlanning #Observability #DevOps #PlatformEngineering #Resilience #OnCall #FexingoBusiness #BusinessPodcast #Technology Keep every episode free: buymeacoffee.com/fexingo© 2026 Fexingo. All rights reserved. 経済学
エピソード
  • How SRE Teams Master Service Dependency Mapping
    2026/09/11
    In this episode, Lucas and Luna explore how modern SRE teams are using service dependency mapping to prevent cascading failures. We dive into a specific case study of a major fintech platform that reduced incident duration by forty percent after implementing real-time topology visualization. The conversation covers the difference between static diagrams and live dependency graphs, the role of distributed tracing in uncovering hidden coupling, and why understanding your blast radius starts with knowing who depends on you. This is Episode 181 of The Site Reliability Podcast. #SiteReliabilityEngineering #ServiceDependencyMapping #FexingoBusiness #BusinessPodcast #TechInfrastructure #CascadingFailures #DistributedTracing #MicroservicesArchitecture #BlastRadiusAnalysis #ProductionEngineering #SystemResilience #CloudComputing #DevOpsCulture #IncidentManagement #TopologicalVisualization #LatencyOptimization #TechLeadership #OperationalExcellence Keep every episode free: buymeacoffee.com/fexingo
    続きを読む 一部表示
    10 分
  • How SRE Teams Master Cross-Team Dependency Mapping
    2026/09/09
    Most outages aren't caused by code bugs but by broken handoffs between teams. In this episode, we examine how top site reliability engineering groups map cross-team dependencies to prevent cascading failures. We look at the specific case of a major cloud provider that reduced incident frequency by forty percent after implementing strict dependency contracts between frontend and backend services. Lucas and Luna break down why organizational boundaries create technical debt, how to identify invisible coupling in microservices architectures, and practical steps for building resilience through better communication protocols rather than just better monitoring tools. #SiteReliabilityEngineering #DependencyMapping #MicroservicesArchitecture #IncidentPrevention #TechnicalDebt #OrganizationalDesign #CascadingFailures #ServiceContracts #CloudInfrastructure #ProductionEngineering #SystemResilience #TeamBoundaries #APIGovernance #SLOManagement #DevOpsCulture #FexingoBusiness #BusinessPodcast #TechLeadership Keep every episode free: buymeacoffee.com/fexingo
    続きを読む 一部表示
    11 分
  • How SRE Teams Manage Technical Debt in Production
    2026/09/08
    Technical debt is often treated as a secondary concern, but for Site Reliability Engineering teams, it is the primary driver of systemic fragility. In this episode, we examine how leading operations groups move beyond simple code refactoring to address architectural and operational debt that accumulates in production environments. We look at specific strategies for identifying invisible liabilities, such as hard-coded dependencies and undocumented manual workarounds, and how teams quantify the cost of inaction using error budgets and toil metrics. By treating technical debt not as a backlog item but as a continuous risk factor, SREs can maintain system stability while still delivering new features. This discussion offers concrete frameworks for prioritizing remediation efforts based on actual impact rather than developer preference. #SiteReliabilityEngineering #TechnicalDebt #ProductionEngineering #SystemFragility #OperationalExcellence #ErrorBudgets #ToilReduction #CloudArchitecture #DevOpsCulture #IncidentPrevention #SystemDesign #TechLeadership #EngineeringMetrics #RiskManagement #ContinuousImprovement #FexingoBusiness #BusinessPodcast #TechTalk Keep every episode free: buymeacoffee.com/fexingo
    続きを読む 一部表示
    9 分
adbl_web_anon_alc_button_suppression_t1
まだレビューはありません