PagerDuty Incident Pod Scheduling — จัดการ

PagerDuty Pod Scheduling

PagerDuty Incident Management Kubernetes Pod Scheduling Alert Prometheus On-call Escalation Auto-remediation Runbook MTTA MTTR
| Pod Status | Cause | PagerDuty Severity | Auto-remediation | SLA |
|---|---|---|---|---|
| Pending (Resource) | CPU/Memory insufficient | High | Cluster Autoscaler | 15 min |
| Pending (Scheduling) | Affinity/Taint mismatch | High | Fix labels or tolerations | 30 min |
| CrashLoopBackOff | App crash, config error | Critical | Rollback deployment | 10 min |
| ImagePullBackOff | Image not found, auth fail | High | Check registry, fix secret | 15 min |
| OOMKilled | Memory limit exceeded | High | Increase memory limit | 15 min |
| Evicted | Node disk pressure | Warning | Cleanup disk, expand PV | 30 min |
เคล็ดลับ
- Noise: ลด Alert Noise ส่งเฉพาะ Actionable Alerts ไม่ส่ง Info
- Runbook: เขียน Runbook ทุก Alert ให้ On-call ทำตามได้ทันที
- Auto: ใช้ Auto-remediation สำหรับปัญหาที่ซ้ำบ่อยและมีวิธีแก้ชัดเจน
- Postmortem: ทำ Blameless Postmortem ทุก Major Incident
- Training: ฝึก On-call ใหม่ Shadow On-call 1-2 สัปดาห์ก่อนเข้า Rotation
การดูแลระบบในสภาพแวดล้อม Production

การบริหารจัดการระบบ Production ที่ดีต้องมี Monitoring ครอบคลุม ใช้เครื่องมืออย่าง Prometheus + Grafana สำหรับ Metrics Collection และ Dashboard หรือ ELK Stack สำหรับ Log Management ตั้ง Alert ให้แจ้งเตือนเมื่อ CPU เกิน 80% RAM ใกล้เต็ม หรือ Disk Usage สูง
Backup Strategy ต้องวางแผนให้ดี ใช้หลัก 3-2-1 คือ มี Backup อย่างน้อย 3 ชุด เก็บใน Storage 2 ประเภทต่างกัน และ 1 ชุดต้องอยู่ Off-site ทดสอบ Restore Backup เป็นประจำ อย่างน้อยเดือนละครั้ง เพราะ Backup ที่ Restore ไม่ได้ก็เหมือนไม่มี Backup
เนื้อหาเกี่ยวข้อง — ทำความเข้าใจ React Suspense MLOps Workflow
เรื่อง Security Hardening ต้องทำตั้งแต่เริ่มต้น ปิด Port ที่ไม่จำเป็น ใช้ SSH Key แทน Password ตั้ง Fail2ban ป้องกัน Brute Force อัพเดท Security Patch สม่ำเสมอ และทำ Vulnerability Scanning อย่างน้อยเดือนละครั้ง ใช้หลัก Principle of Least Privilege ให้สิทธิ์น้อยที่สุดที่จำเป็น
แนะนำเพิ่มเติม — สัญญาณเทรดรายวัน XM Signal
PagerDuty คืออะไร
Incident Management Alert Monitoring Prometheus Datadog On-call SMS Phone Email Escalation Timeline Postmortem Analytics MTTA MTTR
เนื้อหาเกี่ยวข้อง — ทำความเข้าใจ Flux CD GitOps Machine Learning Pipeline
Pod Scheduling Problem คืออะไร
Kubernetes Pod Node Resource CPU Memory Affinity Taint Toleration PV Priority Quota NotReady ImagePullBackOff CrashLoopBackOff OOMKilled
ตั้งค่า Alert อย่างไร
PagerDuty Service Integration Key Alertmanager Prometheus Alert Rules Severity Routing Critical High Low Escalation 5 นาที Repeat
แนะนำเพิ่มเติม — หนังสือเทรดที่ SiamCafeBook
เนื้อหาเกี่ยวข้อง — แนะนำให้อ่าน CDK Construct Container Orchestration
Auto-remediation ทำอย่างไร
Rundeck Script Scale Node Autoscaler Delete Pod ReplicaSet Drain Node Cleanup Images Logs Certificate Renew Runbook On-call ขั้นตอน
สรุป
PagerDuty Incident Pod Scheduling Kubernetes Alert Prometheus Alertmanager On-call Escalation Auto-remediation Runbook MTTA MTTR Postmortem Production
เนื้อหาเกี่ยวข้อง — ทำความเข้าใจ Whisper Speech Progressive Delivery





