L2 Support: diagnose, mitigate, and escalate
Six lessons, 30 questions, and nine cases to practice triage, diagnosis, mitigation, escalation, and L2 support continuity.
Objectives and progression
A professional assessment course with six modules on production-support decisions. Interpret Linux, DNS, HTTP, TLS, PostgreSQL, and Kubernetes signals; practice reconciliation before retry, English communication, handover, and improvement. Includes fictional funds and APS scenarios and an internal assessment of 27 decisions in 60 minutes. L2/L3 boundaries and authorizations depend on the organization.
Audience: APS L2/L3, infrastructure, SRE, systems administration teams, and technical managers.
Prerequisites: Application, monitoring, and production-support fundamentals; no bank-specific internal process assumed.
300 estimated study minutes
- Distinguish symptoms, impact, and urgency to organize the response.
- Build a reproducible diagnosis without exposing sensitive data.
- Interpret service, DNS, HTTP, TLS, and resource states.
- Choose controlled actions and demonstrate service recovery.
- Provide actionable context and keep stakeholders informed.
- Maintain continuity and turn incidents into verifiable actions.
Modules
- Triage based on impact
- Useful hypotheses and evidence
- Diagnosis by layer
- Operational mitigation and validation
- Escalation and communication in English
- Handover and improvement
Continue learning
References and version
Operational support; PostgreSQL 18, OpenSSL 3.5 and BIND 9.20.29 examples; reviewed 2026-09-30
- Effective troubleshooting · 2026-09-30
- Incident response · 2026-09-30
- Postmortem culture · 2026-09-30
- systemctl manual · 2026-09-30
- journalctl manual · 2026-09-30
- df manual · 2026-09-30
- curl command line manual · 2026-09-30
- openssl s_client · 2026-09-30
- BIND dig manual · 2026-09-30
- Monitoring database activity · 2026-09-30
- Viewing locks · 2026-09-30
- Logging cheat sheet · 2026-09-30
- Pod lifecycle · 2026-09-30
- Debug pods · 2026-09-30
- Timeouts retries and backoff with jitter · 2026-09-30
What you will explore
0 / 6Triage based on impact
Distinguish symptoms, impact, and urgency to organize the response.
Useful hypotheses and evidence
Build a reproducible diagnosis without exposing sensitive data.
Diagnosis by layer
Interpret service, DNS, HTTP, TLS, and resource states.
Operational mitigation and validation
Choose controlled actions and demonstrate service recovery.
Escalation and communication in English
Provide actionable context and keep stakeholders informed.
Handover and improvement
Maintain continuity and turn incidents into verifiable actions.