← Back to catalogue
Professional assessment

Production Support L3: investigate and recover

Investigate incidents, recover batch and file flows, and confirm business outcomes. Practice communication, escalation, reconciliation, and bounded automation.

BigSavantAvailable
BigSavantL312 lessons

Objectives and progression

Twelve modules prepare production-support decisions from triage through post-incident improvement. Lessons connect hypotheses, batch, transfers, indicators, handover, and runbooks to fictional financial-operations examples. The bank includes calculations, multiple-response questions, and twenty scenarios with constraints and alternatives. Includes Linux, JVM, TLS, and DNS diagnosis, retries under load, and Kubernetes rollout interpretation, distributed observability, access control, and recovery boundaries. AutoSys and Transfer CFT provide tool context; concrete commands and codes must be confirmed against the installed version. This assesses knowledge and decisions rather than awarding an external certification or demonstrating hands-on production access.

Audience: L2/L3 support, Application Production Support, and technical managers with operating responsibilities.

Prerequisites: Basic systems, applications, logs, and data-flow knowledge. Linux and middleware familiarity is recommended.

460 estimated study minutes

  • Interpret telemetry coverage, percentiles, and recovery with authority and integrity.
  • Distinguish disk, log, resource-pressure, JVM, TLS, and DNS evidence.
  • Control retry amplification and interpret rollout gates and conditions.
  • Connect impact, urgency, and response coordination.
  • Investigate hypotheses and escalate with actionable evidence.
  • Recover batch and files considering partial state and duplicates.
  • Validate outcomes, handover, and preventive-action effectiveness.

Modules

  1. Triage and organize the response
  2. Investigate hypotheses and dependencies
  3. Recover batch chains without falsifying success
  4. Validate transfers and reconciliation
  5. Measure impact and confirm recovery
  6. Turn incidents into verifiable improvement
  7. Shift handover and escalation
  8. Runbooks that support decisions
  9. Linux, JVM, and connectivity diagnosis
  10. Recovery under load and after releases
  11. Observability and distributed evidence
  12. Access, change, and recovery boundaries

Continue learning

Technical assessments

References and version

DR Production Support L3 2026.4; Linux, JDK 25 HotSpot, OpenSSL 3.5 and Kubernetes examples require installed-version checks

What you will explore

0 / 12

Learning is also trying.

Original explained questions, flashcards, and scenarios to apply the concepts.

Practice
Independent professional path with lessons and assessment of the presented topics. It does not award external certification or replace supervised practical experience.

Assessment topics

Incident management—
Investigation and evidence—
Batch and dependencies—
Files and reconciliation—
Measurement and recovery—
Problems and prevention—
Handover and escalation—
Runbooks and automation—
Production technical diagnosis—
Recovery and change control—
Observability and evidence quality—
Operational control and recovery—