PagerDuty Incident Pod Scheduling — จัดการ

PagerDuty Pod Scheduling

PagerDuty Incident Management Kubernetes Pod Scheduling Alert Prometheus On-call Escalation Auto-remediation Runbook MTTA MTTR
| Pod Status | Cause | PagerDuty Severity | Auto-remediation | SLA |
|---|---|---|---|---|
| Pending (Resource) | CPU/Memory insufficient | High | Cluster Autoscaler | 15 min |
| Pending (Scheduling) | Affinity/Taint mismatch | High | Fix labels or tolerations | 30 min |
| CrashLoopBackOff | App crash, config error | Critical | Rollback deployment | 10 min |
| ImagePullBackOff | Image not found, auth fail | High | Check registry, fix secret | 15 min |
| OOMKilled | Memory limit exceeded | High | Increase memory limit | 15 min |
| Evicted | Node disk pressure | Warning | Cleanup disk, expand PV | 30 min |
เคล็ดลับ
- Noise: ลด Alert Noise ส่งเฉพาะ Actionable Alerts ไม่ส่ง Info
- Runbook: เขียน Runbook ทุก Alert ให้ On-call ทำตามได้ทันที
- Auto: ใช้ Auto-remediation สำหรับปัญหาที่ซ้ำบ่อยและมีวิธีแก้ชัดเจน
- Postmortem: ทำ Blameless Postmortem ทุก Major Incident
- Training: ฝึก On-call ใหม่ Shadow On-call 1-2 สัปดาห์ก่อนเข้า Rotation
การดูแลระบบในสภาพแวดล้อม Production

การบริหารจัดการระบบ Production ที่ดีต้องมี Monitoring ครอบคลุม ใช้เครื่องมืออย่าง Prometheus + Grafana สำหรับ Metrics Collection และ Dashboard หรือ ELK Stack สำหรับ Log Management ตั้ง Alert ให้แจ้งเตือนเมื่อ CPU เกิน 80% RAM ใกล้เต็ม หรือ Disk Usage สูง
Backup Strategy ต้องวางแผนให้ดี ใช้หลัก 3-2-1 คือ มี Backup อย่างน้อย 3 ชุด เก็บใน Storage 2 ประเภทต่างกัน และ 1 ชุดต้องอยู่ Off-site ทดสอบ Restore Backup เป็นประจำ อย่างน้อยเดือนละครั้ง เพราะ Backup ที่ Restore ไม่ได้ก็เหมือนไม่มี Backup
เรื่อง Security Hardening ต้องทำตั้งแต่เริ่มต้น ปิด Port ที่ไม่จำเป็น ใช้ SSH Key แทน Password ตั้ง Fail2ban ป้องกัน Brute Force อัพเดท Security Patch สม่ำเสมอ และทำ Vulnerability Scanning อย่างน้อยเดือนละครั้ง ใช้หลัก Principle of Least Privilege ให้สิทธิ์น้อยที่สุดที่จำเป็น
PagerDuty คืออะไร
Incident Management Alert Monitoring Prometheus Datadog On-call SMS Phone Email Escalation Timeline Postmortem Analytics MTTA MTTR
Pod Scheduling Problem คืออะไร
Kubernetes Pod Node Resource CPU Memory Affinity Taint Toleration PV Priority Quota NotReady ImagePullBackOff CrashLoopBackOff OOMKilled
ตั้งค่า Alert อย่างไร
PagerDuty Service Integration Key Alertmanager Prometheus Alert Rules Severity Routing Critical High Low Escalation 5 นาที Repeat
Auto-remediation ทำอย่างไร
Rundeck Script Scale Node Autoscaler Delete Pod ReplicaSet Drain Node Cleanup Images Logs Certificate Renew Runbook On-call ขั้นตอน
สรุป
PagerDuty Incident Pod Scheduling Kubernetes Alert Prometheus Alertmanager On-call Escalation Auto-remediation Runbook MTTA MTTR Postmortem Production





