『DevOps Daily with Fexingo: CI/CD, Kubernetes, and Modern Software Operations』のカバーアート

DevOps Daily with Fexingo: CI/CD, Kubernetes, and Modern Software Operations

DevOps Daily with Fexingo: CI/CD, Kubernetes, and Modern Software Operations

著者: Fexingo
無料で聴く

Lucas and Luna dissect the daily realities of DevOps, from CI/CD pipeline design to Kubernetes cluster management and the human systems that keep software running. Each episode grounds abstract principles in real incidents—a failed deployment at a major retailer, a postmortem from a cloud outage, a configuration drift disaster—and traces the operational decisions that turned them around. Lucas brings the technical precision of a working engineer, while Luna pushes on the team dynamics, cost trade-offs, and organizational bottlenecks that separate resilient operations from fragile ones. They discuss monitoring strategies, incident response playbooks, infrastructure-as-code trade-offs, and the cultural friction between development velocity and operational stability—always with concrete examples, never with buzzwords. This is the show for engineers, SREs, and platform leads who want to hear two seasoned practitioners argue through the hard choices: when to rewrite vs. patch, how much observability is enough, and how to keep a multi-cloud deployment from becoming a management nightmare. By the end, you'll carry away a sharpened question about your own stack and a new way to think about reliability. #DevOps #CICD #Kubernetes #SiteReliabilityEngineering #PipelineAutomation #InfrastructureAsCode #IncidentResponse #Monitoring #Observability #CloudOperations #ContainerOrchestration #Postmortem #DeploymentStrategy #Technology #FexingoBusiness #BusinessPodcast #SoftwareEngineering #PlatformEngineering Keep every episode free: buymeacoffee.com/fexingo© 2026 Fexingo. All rights reserved. 経済学
エピソード
  • How Kubernetes Watch Events Overload the API Server
    2026/07/21
    In this episode of DevOps Daily, Lucas and Luna dive into a notoriously overlooked Kubernetes failure mode—API server overload caused by excessive watch events. They break down how controllers, informers, and custom operators all lean on watch-based architectures, and why a single misconfigured watch can cascade into cluster-wide latency or even outages. Through the lens of a real incident at a mid-sized fintech, they explain what happened when a poorly tuned deployment triggered millions of watch events per minute, flooded etcd, and brought the control plane to its knees. Lucas outlines the specific metrics to monitor (watch event rate, request latency, etcd db size) and the architectural patterns—like using DeltaFIFO, filtering with field selectors, and throttling informer resync intervals—that prevent these meltdowns. Luna brings her own war story about a CI/CD operator that hammered the API server during a Helm release, and the two discuss how to audit your cluster for watch-heavy workloads. By the end, listeners will know exactly how to diagnose and prevent one of the most common yet invisible causes of Kubernetes instability. #Kubernetes #APIServer #WatchEvents #Etcd #K8sOperators #InformerPattern #ClusterReliability #DevOps #CI/CD #Fintech #SiteReliabilityEngineering #K8sPerformance #ControlPlane #DeltaFIFO #FieldSelectors #ContainerOrchestration #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo
    続きを読む 一部表示
    10 分
  • How Kubernetes Vertical Pod Autoscaler Recommends Wrong Requests
    2026/07/21
    Kubernetes Vertical Pod Autoscaler (VPA) is supposed to right-size container resource requests automatically. But in practice, VPA's recommender often suggests CPU and memory values that lead to pod evictions or wasted capacity. Drawing on real incidents from a mid-2026 production cluster running 50 microservices, Lucas and Luna unpack why VPA's default OOM-aware policy can misinterpret short-lived memory spikes, how its recommendation window of 8 days lags behind traffic patterns, and why the 'lower bound' mode sometimes recommends requests below actual usage. They walk through a specific case where VPA recommended 512 MiB of memory for a Go service that actually needed 1.2 GiB under peak load, causing repeated OOMKills. The episode closes with practical mitigations: setting custom OOM-scoring thresholds, using VPA in 'initial' mode for batch workloads, and combining VPA with Horizontal Pod Autoscaler using resource metrics. Listeners come away with a clear mental model of VPA's blind spots and how to compensate for them. #Kubernetes #VPA #VerticalPodAutoscaler #ResourceRequests #OOMKill #K8sAutoscaling #ContainerSizing #GoService #PodEviction #HPA #ResourceMetrics #ClusterOps #PerformanceTuning #DevOps #Technology #FexingoBusiness #BusinessPodcast #DevOpsDaily Keep every episode free: buymeacoffee.com/fexingo
    続きを読む 一部表示
    10 分
  • How Kubernetes StatefulSet Rolling Update Breaks Persistent Volume Claims
    2026/07/20
    In this episode of DevOps Daily, Lucas and Luna dive into a subtle but destructive Kubernetes pitfall: StatefulSet rolling updates that orphan PersistentVolumeClaims when the pod template changes. Using a real-world incident at a mid-sized fintech, they walk through how a simple image update triggered PVC detachment, stalled the rollout, and left the cluster in an inconsistent state. They explain the underlying mechanics — how StatefulSet's pod identity logic interacts with PVC templates, why the 'OnDelete' update strategy buys time but doesn't fix the root cause, and what monitoring signals to watch for. Lucas shares a concrete remediation: pre-binding PVCs via volumeClaimTemplates with stable names, and using a custom operator to handle canary updates safely. Listeners will learn one specific configuration pattern to prevent this silent failure, plus a diagnostic command to check for orphaned PVCs after any StatefulSet change. This is a focused, practical episode for anyone running stateful workloads on Kubernetes. #Kubernetes #StatefulSet #PersistentVolumeClaim #RollingUpdate #DevOps #CloudNative #StatefulWorkloads #K8sFailure #PodIdentity #VolumeClaimTemplates #OnDeleteStrategy #Fintech #IncidentResponse #KubernetesTroubleshooting #CI/CD #Technology #FexingoBusiness #DevOpsPodcast Keep every episode free: buymeacoffee.com/fexingo
    続きを読む 一部表示
    9 分
adbl_web_anon_alc_button_suppression_t1
まだレビューはありません