Production Support L3: investigate and recover
Investigate incidents, recover batch and file flows, and confirm business outcomes. Practice communication, escalation, reconciliation, and bounded automation.
Objectives and progression
Twelve modules prepare production-support decisions from triage through post-incident improvement. Lessons connect hypotheses, batch, transfers, indicators, handover, and runbooks to fictional financial-operations examples. The bank includes calculations, multiple-response questions, and twenty scenarios with constraints and alternatives. Includes Linux, JVM, TLS, and DNS diagnosis, retries under load, and Kubernetes rollout interpretation, distributed observability, access control, and recovery boundaries. AutoSys and Transfer CFT provide tool context; concrete commands and codes must be confirmed against the installed version. This assesses knowledge and decisions rather than awarding an external certification or demonstrating hands-on production access.
Audience: L2/L3 support, Application Production Support, and technical managers with operating responsibilities.
Prerequisites: Basic systems, applications, logs, and data-flow knowledge. Linux and middleware familiarity is recommended.
460 estimated study minutes
- Interpret telemetry coverage, percentiles, and recovery with authority and integrity.
- Distinguish disk, log, resource-pressure, JVM, TLS, and DNS evidence.
- Control retry amplification and interpret rollout gates and conditions.
- Connect impact, urgency, and response coordination.
- Investigate hypotheses and escalate with actionable evidence.
- Recover batch and files considering partial state and duplicates.
- Validate outcomes, handover, and preventive-action effectiveness.
Modules
- Triage and organize the response
- Investigate hypotheses and dependencies
- Recover batch chains without falsifying success
- Validate transfers and reconciliation
- Measure impact and confirm recovery
- Turn incidents into verifiable improvement
- Shift handover and escalation
- Runbooks that support decisions
- Linux, JVM, and connectivity diagnosis
- Recovery under load and after releases
- Observability and distributed evidence
- Access, change, and recovery boundaries
Continue learning
References and version
DR Production Support L3 2026.4; Linux, JDK 25 HotSpot, OpenSSL 3.5 and Kubernetes examples require installed-version checks
- Incident response · 2026-09-29
- Managing incidents · 2026-09-29
- Effective troubleshooting · 2026-09-29
- Monitoring distributed systems · 2026-09-29
- Implementing SLOs · 2026-09-29
- Alerting on SLOs · 2026-09-29
- Postmortem culture · 2026-09-29
- Postmortem practices · 2026-09-29
- Evolution of automation · 2026-09-29
- Being on-call · 2026-09-29
- AutoSys: success_codes and exit status · 2026-09-29
- Transfer CFT product overview · 2026-09-29
- Making retries safe with idempotent APIs · 2026-09-29
- Data processing pipelines · 2026-09-29
- Data integrity · 2026-09-29
- Process file descriptors · 2026-10-01
- Filesystem byte and inode usage · 2026-10-01
- Unlink and open file lifetime · 2026-10-01
- Journal selection and time display · 2026-10-01
- Pressure Stall Information · 2026-10-01
- Java diagnostic tools · 2026-10-01
- OpenSSL client diagnostics · 2026-10-01
- DNS resolution debugging · 2026-10-01
- Timeouts, retries, backoff and jitter · 2026-10-01
- Handling overload · 2026-10-01
- Deployment progress and availability · 2026-10-01
- Canary release evaluation · 2026-10-01
- Context propagation · 2026-10-01
- Sampling · 2026-10-01
- Query functions · 2026-10-01
- Histograms and summaries · 2026-10-01
- Metric and label naming · 2026-10-01
- Root cause analysis concepts · 2026-10-01
- Privileged Identity Management overview · 2026-10-01
- Handling sensitive data · 2026-10-01
- Configuring fencing in a Red Hat High Availability cluster · 2026-10-01
- Plan for Disaster Recovery · 2026-10-01
What you will explore
0 / 12Triage and organize the response
Connect symptoms to business impact and separate coordination, communication, and execution.
Investigate hypotheses and dependencies
Build a discriminating diagnosis across Linux, JVM, networking, and databases.
Recover batch chains without falsifying success
Connect the scheduler, exit codes, calendars, and result validation.
Validate transfers and reconciliation
Separate transport, receipt, processing, and functional acceptance.
Measure impact and confirm recovery
Use service indicators, objectives, and observation windows to assess stability.
Turn incidents into verifiable improvement
Build a blameless analysis and actions with owners, deadlines, and effectiveness criteria.
Shift handover and escalation
Transfer context and ownership without losing continuity.
Runbooks that support decisions
Design steps with preconditions, limits, and recovery evidence.
Linux, JVM, and connectivity diagnosis
Choose observations distinguishing resource shortage, application waiting, and connection failures.
Recovery under load and after releases
Control retries, capacity, and rollout before declaring functional recovery.
Observability and distributed evidence
Interpret traces, metrics, and problems without confusing observation gaps with business outcomes.
Access, change, and recovery boundaries
Plan interventions with authority, compatible data, and verifiable recovery criteria.