← All posts
Operations ·

L1, L2 and L3 Support: ITIL in Operations

What ITIL actually says about tiered support, how to build a team around it, and a case where all three tiers existed and still did not hold.

ITIL carries support on three practices: the service desk, incident management and problem management. In operations those practices are usually distributed across the teams through a tier model (L1, L2 & L3). A clean assignment of the work builds a team that holds up when it matters.

What ITIL means by L1, L2 and L3

The service desk captures demand for incident resolution and service requests and is the single point of contact between provider and users. Incident management minimizes the negative impact of incidents by restoring normal service operation as quickly as possible. Problem management reduces the likelihood and impact of future incidents by identifying causes and managing workarounds and known errors.

Terms per ITIL
Incident

An unplanned interruption to a service, or a reduction in the quality of a service.

Problem

A cause, or potential cause, of one or more incidents.

Known error

A problem that has been analysed but has not been resolved.

Workaround

A solution that reduces or eliminates the impact while a full resolution is missing.

ITIL defines escalation as sharing awareness or transferring ownership of an issue or work item. Those are two different acts inside one definition. Informing someone is not a transfer of ownership. Skip that distinction and you end up with a case three people know about and nobody owns.

Tier Practice in ITIL Owns Hands over when
L1 Service desk Intake, prioritization, communication, resolution by runbook the case falls outside the runbook
L2 Incident management, support team Diagnosis, workaround, fix, runbook upkeep the cause recurs or touches the design
L3 Problem management, technical practices Root cause, known errors, architecture, standards the fix needs a change through change control

A support team focuses on maintaining normal operations and resolves user requests, incidents and problems for specified products or services. Routing goes by the category of the incident, not by the experience of whoever happens to be reachable.

An important part that tends to get overlooked is the major incident process. A major incident is an incident with significant business impact requiring an immediate coordinated resolution. ITIL names swarming for this. Many stakeholders work together initially, until it is clear who continues and who moves on.

Report: alert, ticket, request L1 · SERVICE DESK Intake, triage, communication Resolves with runbook and known error Stays the owner of the ticket escalates beyond the runbook L2 · SUPPORT TEAM Diagnosis, workaround, fix Takes what the runbook does not cover Maintains runbooks and automation escalates on a recurring cause L3 · PROBLEM MANAGEMENT Cause, known error, design Works on the repetition, not the case Off the on-call rotation Known error and runbook back to L1 MAJOR INCIDENT All tiers at the same time One person leads and decides No tier path, no handover

Ownership

  1. ShapeL1 broad and close to the users. L2 cut by product or platform, not by technology. L3 small and off the rotation by default. If L3 gets called every week, that is not a staffing question, it is a gap in the runbook.
  2. OwnershipThe ticket stays with L1, even when the technical work moves. L1 runs communication to the user until closure. That keeps escalation a handover of work instead of the disappearance of the case.
  3. KnowledgeA runbook names the dashboard, the log line, the command, and the check that it worked. ITIL puts it more generally: knowledge is information in the context of whoever needs it. A 300-page manual does not help at the service desk when an answer is due in two minutes.
  4. Escalation on the clockThe condition for handover lives in the category and the timer, not in the gut feeling of the person on the ticket. Escalating only once you give up means escalating too late.

What that looks like in practice depends on the application landscape, not on a formula. As a starting point this has held up: L1 carries the surface and is staffed continuously. L2 runs at least two people per platform, so that on-call and day work do not block each other. L3 is a small group working to a plan. The question is not how many tiers an operation has, but whether every tier holds the information it needs to work the case without friction.

The interaction needs three fixed points on the agenda, otherwise it falls apart during the incident.

The shift handover clears open cases, running workarounds, and what may fire in the next hours. A weekly review lets L1 report the patterns that stay invisible ticket by ticket. Out of that practice grows a continual improvement process. The post-mortem after every larger case closes the loop by putting its result into the runbook, and therefore back with L1.

Between L1 and L2, a named person beats a group. One member of L2 is the contact for L1 for a week, takes questions and escalations, and records at the end of the week what belongs in the runbook.

For a major incident, three roles have to be named before it happens: one person who leads and decides, one who communicates outward, and the specialists who work. Whoever leads does not work the case. In quiet times that separation looks excessive. During the incident it is the difference between coordination and several people working in parallel.

A task belongs to L1 once it is described, verified and bounded. Described means a runbook. Verified means run through together once. Bounded means a permission that covers exactly this case and nothing beyond it.

What gets measured is not the tier but the path. Four numbers are enough to start: share of cases that end at L1. Time to escalate. Share of recurring causes. Reopen rate.

Careful with metrics

ITIL describes the watermelon effect. An SLA is green on the outside and red on the inside. Availability reads 99.6 per cent, and the missing 0.4 per cent lands exactly on the business process that mattered. Measure system values only and you get a team hitting its targets and users reporting something else.

The tiers were there, the handovers were not…

Starting point

In an operation I took over, L1, L2 and L3 already existed. On paper the model was complete. Day to day it was sluggish.

Communication between the tiers was mediocre. Information was passed along, but not handed over. A case moved a tier and came back with the same questions somebody had already asked. Part of the work was declared nowhere. Nobody could say whether a particular check belonged to L1 or to L2. So sometimes one side did it, sometimes the other, sometimes neither.

The gap showed most clearly in a crisis. It was not defined who takes the lead. Several capable people worked the same case and no decisions were made, because everyone assumed that was not their role. The time did not go into analysis. It went into agreeing who analyses.

The finding

The structure was not the problem. The conditions were.

L1 had neither the permissions nor the tooling for tasks it could have handled without difficulty. So practically everything travelled upward. And what arrived upward arrived without context, because there was no form in which context travels.

Tooling and permissions

I evaluated the existing tools against a single criterion: what actually connects the tiers? A shared view of the tickets, access to the same dashboards and logs, traceable actions. Not one tool per team, but one path through the teams.

Then I enabled the teams across the boundary. L1 got access, permissions and instructions for a defined set of tasks that had sat with L2. Every one of those tasks came with three conditions: a runbook that covers the case, a permission scoped exactly that far, and a defined way back when it does not work. Without those three, a moved task becomes a new source of failure.

Three conditions
Runbook

Names the dashboard, the command and the check that it worked.

Permission

Reaches exactly as far as the task and no further.

Way back

Says who takes over when the runbook does not hold.

The side effect was the real gain. Once L1 could close a case itself, passing it on stopped being a reflex. People read a case differently when they are the ones who resolve it.

Post-mortems

The second step was the post-mortem. Before, a case ended with restoration. After, it ended with a review: what happened, when it was noticed, what stretched it out, which action follows, and who owns that action.

The effect was larger than expected, and it was not technical. The post-mortem built trust, not only on the IT side but toward the business customer. It was run without blame, so details came out that had gone unmentioned before. With the details the quality of the analysis rose, and with it the quality of the actions.

Out of the trust came further openings. Topics nobody would have raised before turned into entries in the improvement backlog. For the first time, operations had a list that came out of reality rather than out of a planning round.

I did not introduce the post-mortem as a rule but on actual cases. I facilitated and wrote up the first few myself. Only after that did it become a fixed part of the work. A template simplifies the process and keeps the integrity and quality of the post-mortem intact, independent of the incident.

Possible fields: timeline, impact, immediate cause, contributing factors, actions with owner and date, and the entry stating what of it goes back into the runbook. Alongside it, a structure for the handover to L1 and back.

POST-MORTEM · TEMPLATE 01 · TIMELINE What happened when, and when it was noticed 02 · IMPACT Who was affected, and for how long 03 · IMMEDIATE CAUSE What triggered the outage 04 · CONTRIBUTING FACTORS What stretched it out 05 · ACTIONS Each with an owner and a date 06 · WAY BACK What of it returns to the runbook Back to L1 Runbook, known error, permission

A template does not take the thinking off anyone. It takes away the decision about what the document looks like, and that decision costs the same quarter of an hour on every case. After a handful of cases the write-ups were comparable. That made patterns visible which ran across several cases and would never have surfaced in a single report.

What did not work right away

Two things took longer than planned.

The first was the fear of breaking something. Permissions alone change no behaviour. The first cases had to be accompanied until L1 accepted the new task as its own. The second was the temptation to improve everything at once. The improvements that came out of a post-mortem for L2, in the form of new runbooks, were sluggish at first and did not get implemented in time. What helped was cutting the leftovers into small packages and tasks, and rolling the improvements out incrementally.

What changed

The information flow was the biggest difference. A case carried its context because there was a form for it. The work was declared, so the negotiation about who is responsible fell away. Escalation paths got shorter because a share of the cases stopped changing tiers at all. And escalation took less time, because a handover had a form instead of a follow-up question.

BEFORE Report L1 L2 L3 a follow-up question at every handover Work declared nowhere, knowledge scattered L1 without permissions or tooling No named lead in a crisis The case changes tier without its context AFTER Report L1 L2 L3 Post-mortem: runbook, template, permission L1 resolves with permission, tool, runbook Escalation only when it is needed Named lead in a major incident Every larger case ends in a post-mortem
What I take from it
  • Tiers without permissions produce queues, not relief.
  • The lead in a crisis has to be named before the crisis starts.
  • A post-mortem works on trust first and on technology second.
  • A template is not paperwork, it is decision time saved on every case.
  • Whatever is not declared ends up being done by nobody.

Where to start

Two questions are enough for the diagnosis. First: which tasks could the tiers take on today if the permission and the runbook existed? Second: who leads the next major incident, and does that person know it?

If either answer is unclear, that is where the work is. If both are unclear, the order stays the same: ownership first, tooling second.

This is the work I do with operations teams. If that sounds like your current situation, get in touch.

Work together?

I help teams turn ad-hoc operations into resilient, automated systems.

Request a call