Feature Tutorials
Feature Tutorials
Welcome to Tutorials. Here you will find information and how-tos for PagerDuty features, organized by performance or improvement goal.
PagerDuty has a huge library of resources available for you to learn more about our products to help you improve your incident response and service reliability. Here you can find collections of resources focused on performance and improvement goals, gathered from our various resources. So whether you prefer to learn by reading, watching, or coding, we've got you covered.
Reduce Mean Time to Resolve (MTTR)
Mean Time to Resolve, or MTTR, can be a useful measure for tracking your team's improvement in incident response.
Learn more about MTTR
- Incident Workflows
- Event Orchestration
- Incident Priority
- Coming soon:
- Probable Origin — ML-suggested likely root cause on an incident
- Related Incidents / Past Incidents — surfaces similar prior incidents for faster triage
- Outlier Incident — flags incidents that deviate from a service's normal pattern
- Operations Console — centralized live view for triaging across many incidents at once
- Incident Workflows — automated, repeatable response sequences triggered on incident creation
Already have MTTR under control and want to go deeper? Check out this podcast episode.
Prevent Burnout & Alert Fatigue
Alert Fatigue can hit any team or any responder. The volume of alerts that many teams receive makes it hard to continue to make decisions about what is happening, whether an alert is meaningful or not. Manage your alerts better by allowing PagerDuty to help. Learn more about Alert Fatigue.
- Basics of Alert Grouping
- Content-Based Alert Grouping
- Intelligent Alert Grouping
- Global Alert Grouping
- Coming Soon:
- Auto-Pause Incident Notifications — suppresses alerts likely to self-resolve
- Round Robin Scheduling — distributes on-call load evenly across a team
- Shift-Based Schedules — modern scheduling model (as distinct from legacy Schedule Basics)
- Dynamic Notifications — adjusts notification urgency based on alert payload
Alert fatigue has origins in neuroscience. Learn more on our podcast.
Improve System Reliability:
Complex systems are complex. Knowing what is going on in your ecosystem is important when trying to keep your systems reliable and provide good user experience.
- Keep track of which services have been changed using Change Events
- Proactively manage alerts during planned work with Maintenance Windows
- Use Incident Priority to communicate the importance of particular incidents
- Coming Soon
- Service Standards — tracks whether services meet configuration best practices
- Operational Reviews — scheduled, data-driven reviews of incident impact over time
- Service Performance Insights — analytics on a service's incident trends
Streamline Incident Communication:
Keeping all stakeholders informed during an incident can be challenging - you have more information to give to internal stakeholders than you do to customers or other external stakeholders. Automating your incident communications helps keep everyone up to date on the status of a live incident.
- General information about Stakeholder Communications in our knowledge base: Managing subscribers on an incident, sending status updates, status update notifcations methods.
- Internal Stakeholder Updates via Slack using Incident Workflows
- Status Pages provide a one-stop shop for folks who need to know more about the status of your services, whether internal or external
- Coming Soon
- Status Update Templates — pre-built templates for consistent stakeholder updates
- Conference Bridge — auto-attach a call bridge to incidents
- Slack / Microsoft Teams ChatOps — manage incidents without leaving chat
Automate Manual Tasks:
Automation is your friend in managing realtime work, avoiding waking human responders after hours, and for increasing speed and reliability in everyday tasks for your team.
- Runbook Automation and Rundeck Open Source is a platform for managing and running automation tools - add your existing scripts and tools and provide users with a secure, auditable platform for automated tasks.
- Automation Actions brings Runbook Automation into your PagerDuty platform, powering autoremediation solutions for your teams.
- Coming Soon
- Custom Incident Actions — one-off webhook-triggered buttons on an incident
- Webhooks — general-purpose outbound event notifications
- Incident Workflows
Not sure what we mean by automation or autoremediation? Check out our Ops Guide.
Gain Visibility into Infrastructure:
Your users only know the "front door" of your services and applications, but your architecture has plenty of hidden complexity. Make sure all stakeholders can be informed, even if they aren't familiar with the services.
- Service Graph and Service Dependencies allows your team to document the relationships and dependencies among your services, illuminating which downstream services are potentially impacted by an ongoing incident.
- Status Pages give you a centralized location to publish incident updates to your interal and external stakeholders.
- Business Service Subscriptions are an easy way for stakeholders to stay aware of user-facing services impacted by incidents in the backend layers of your ecosystem.
- Coming Soon
- Analytics Dashboard / Insights suite (Incident Activity, Responder, Team, Escalation Policy, Business Impact Insights)
- Visibility Console
- Audit Trail Reporting
- Incident Activity Insights
- Service Performance Insights
- Business Impact Insights
- Audit Trail Reporting
- Event Analytics
Maintaining good service ownership hygeine will make finding responders much easier during an incident, but it can be complex to figure out who should own what services. For more on Full Service Ownership, check out our Ops Guide
Team and On-Call Management
Don't just cover the basics; help your team have a better on-call experience. Set better team norms and expectations, manage team health, and make better use of PagerDuty's essential features.
- Coming SoonEscalation Policies (deferred per your earlier note, but the natural home for it)
- Shift-Based Schedules / Round Robin Scheduling
- Teams, Advanced Permissions, User Roles
- Responder Insights — analytics on individual responder load and activity
- Team Insights — analytics on team-level incident load and performance
- Escalation Policy Insights — analytics on how well escalation policies are performing
- On-Call Readiness Reports — surfaces gaps in on-call coverage before they become a problem
- User Onboarding Report — tracks new user activation and setup completion
- Operational Maturity — benchmarks account configuration against best practices
AI and Assistive Agents
PagerDuty AI's agent lineup — Scribe Agent, Shift Agent, Insights Agent, and Paige, PagerDuty's SRE Agent — are functional agents built into the platform, each doing one job, with Paige orchestrating across them. Together with the PagerDuty MCP Server, they go beyond a standard features list.
- Coming Soon
- Scribe Agent
- Shift Agent
- Insights Agent
- Paige
- PagerDuty MCP Server