L1, L2 and L3 Support: ITIL in Operations
What ITIL actually says about tiered support, how to build a team around it, and a case where all three tiers existed and still did not hold.
ITIL carries support on three practices: the service desk, incident management and problem management. In operations those practices are usually distributed across the teams through a tier model (L1, L2 & L3). A clean assignment of the work builds a team that holds up when it matters.
What ITIL means by L1, L2 and L3
The service desk captures demand for incident resolution and service requests and is the single point of contact between provider and users. Incident management minimizes the negative impact of incidents by restoring normal service operation as quickly as possible. Problem management reduces the likelihood and impact of future incidents by identifying causes and managing workarounds and known errors.
An unplanned interruption to a service, or a reduction in the quality of a service.
A cause, or potential cause, of one or more incidents.
A problem that has been analysed but has not been resolved.
A solution that reduces or eliminates the impact while a full resolution is missing.
ITIL defines escalation as sharing awareness or transferring ownership of an issue or work item. Those are two different acts inside one definition. Informing someone is not a transfer of ownership. Skip that distinction and you end up with a case three people know about and nobody owns.
| Tier | Practice in ITIL | Owns | Hands over when |
|---|---|---|---|
| L1 | Service desk | Intake, prioritization, communication, resolution by runbook | the case falls outside the runbook |
| L2 | Incident management, support team | Diagnosis, workaround, fix, runbook upkeep | the cause recurs or touches the design |
| L3 | Problem management, technical practices | Root cause, known errors, architecture, standards | the fix needs a change through change control |
A support team focuses on maintaining normal operations and resolves user requests, incidents and problems for specified products or services. Routing goes by the category of the incident, not by the experience of whoever happens to be reachable.
An important part that tends to get overlooked is the major incident process. A major incident is an incident with significant business impact requiring an immediate coordinated resolution. ITIL names swarming for this. Many stakeholders work together initially, until it is clear who continues and who moves on.
Ownership
- ShapeL1 broad and close to the users. L2 cut by product or platform, not by technology. L3 small and off the rotation by default. If L3 gets called every week, that is not a staffing question, it is a gap in the runbook.
- OwnershipThe ticket stays with L1, even when the technical work moves. L1 runs communication to the user until closure. That keeps escalation a handover of work instead of the disappearance of the case.
- KnowledgeA runbook names the dashboard, the log line, the command, and the check that it worked. ITIL puts it more generally: knowledge is information in the context of whoever needs it. A 300-page manual does not help at the service desk when an answer is due in two minutes.
- Escalation on the clockThe condition for handover lives in the category and the timer, not in the gut feeling of the person on the ticket. Escalating only once you give up means escalating too late.
What that looks like in practice depends on the application landscape, not on a formula. As a starting point this has held up: L1 carries the surface and is staffed continuously. L2 runs at least two people per platform, so that on-call and day work do not block each other. L3 is a small group working to a plan. The question is not how many tiers an operation has, but whether every tier holds the information it needs to work the case without friction.
The interaction needs three fixed points on the agenda, otherwise it falls apart during the incident.
The shift handover clears open cases, running workarounds, and what may fire in the next hours. A weekly review lets L1 report the patterns that stay invisible ticket by ticket. Out of that practice grows a continual improvement process. The post-mortem after every larger case closes the loop by putting its result into the runbook, and therefore back with L1.
Between L1 and L2, a named person beats a group. One member of L2 is the contact for L1 for a week, takes questions and escalations, and records at the end of the week what belongs in the runbook.
For a major incident, three roles have to be named before it happens: one person who leads and decides, one who communicates outward, and the specialists who work. Whoever leads does not work the case. In quiet times that separation looks excessive. During the incident it is the difference between coordination and several people working in parallel.
A task belongs to L1 once it is described, verified and bounded. Described means a runbook. Verified means run through together once. Bounded means a permission that covers exactly this case and nothing beyond it.
What gets measured is not the tier but the path. Four numbers are enough to start: share of cases that end at L1. Time to escalate. Share of recurring causes. Reopen rate.
ITIL describes the watermelon effect. An SLA is green on the outside and red on the inside. Availability reads 99.6 per cent, and the missing 0.4 per cent lands exactly on the business process that mattered. Measure system values only and you get a team hitting its targets and users reporting something else.
The tiers were there, the handovers were not…
Starting point
In an operation I took over, L1, L2 and L3 already existed. On paper the model was complete. Day to day it was sluggish.
Communication between the tiers was mediocre. Information was passed along, but not handed over. A case moved a tier and came back with the same questions somebody had already asked. Part of the work was declared nowhere. Nobody could say whether a particular check belonged to L1 or to L2. So sometimes one side did it, sometimes the other, sometimes neither.
The gap showed most clearly in a crisis. It was not defined who takes the lead. Several capable people worked the same case and no decisions were made, because everyone assumed that was not their role. The time did not go into analysis. It went into agreeing who analyses.
The finding
The structure was not the problem. The conditions were.
L1 had neither the permissions nor the tooling for tasks it could have handled without difficulty. So practically everything travelled upward. And what arrived upward arrived without context, because there was no form in which context travels.
Tooling and permissions
I evaluated the existing tools against a single criterion: what actually connects the tiers? A shared view of the tickets, access to the same dashboards and logs, traceable actions. Not one tool per team, but one path through the teams.
Then I enabled the teams across the boundary. L1 got access, permissions and instructions for a defined set of tasks that had sat with L2. Every one of those tasks came with three conditions: a runbook that covers the case, a permission scoped exactly that far, and a defined way back when it does not work. Without those three, a moved task becomes a new source of failure.
Names the dashboard, the command and the check that it worked.
Reaches exactly as far as the task and no further.
Says who takes over when the runbook does not hold.
The side effect was the real gain. Once L1 could close a case itself, passing it on stopped being a reflex. People read a case differently when they are the ones who resolve it.
Post-mortems
The second step was the post-mortem. Before, a case ended with restoration. After, it ended with a review: what happened, when it was noticed, what stretched it out, which action follows, and who owns that action.
The effect was larger than expected, and it was not technical. The post-mortem built trust, not only on the IT side but toward the business customer. It was run without blame, so details came out that had gone unmentioned before. With the details the quality of the analysis rose, and with it the quality of the actions.
Out of the trust came further openings. Topics nobody would have raised before turned into entries in the improvement backlog. For the first time, operations had a list that came out of reality rather than out of a planning round.
I did not introduce the post-mortem as a rule but on actual cases. I facilitated and wrote up the first few myself. Only after that did it become a fixed part of the work. A template simplifies the process and keeps the integrity and quality of the post-mortem intact, independent of the incident.
Possible fields: timeline, impact, immediate cause, contributing factors, actions with owner and date, and the entry stating what of it goes back into the runbook. Alongside it, a structure for the handover to L1 and back.
A template does not take the thinking off anyone. It takes away the decision about what the document looks like, and that decision costs the same quarter of an hour on every case. After a handful of cases the write-ups were comparable. That made patterns visible which ran across several cases and would never have surfaced in a single report.
What did not work right away
Two things took longer than planned.
The first was the fear of breaking something. Permissions alone change no behaviour. The first cases had to be accompanied until L1 accepted the new task as its own. The second was the temptation to improve everything at once. The improvements that came out of a post-mortem for L2, in the form of new runbooks, were sluggish at first and did not get implemented in time. What helped was cutting the leftovers into small packages and tasks, and rolling the improvements out incrementally.
What changed
The information flow was the biggest difference. A case carried its context because there was a form for it. The work was declared, so the negotiation about who is responsible fell away. Escalation paths got shorter because a share of the cases stopped changing tiers at all. And escalation took less time, because a handover had a form instead of a follow-up question.
- Tiers without permissions produce queues, not relief.
- The lead in a crisis has to be named before the crisis starts.
- A post-mortem works on trust first and on technology second.
- A template is not paperwork, it is decision time saved on every case.
- Whatever is not declared ends up being done by nobody.
Where to start
Two questions are enough for the diagnosis. First: which tasks could the tiers take on today if the permission and the runbook existed? Second: who leads the next major incident, and does that person know it?
If either answer is unclear, that is where the work is. If both are unclear, the order stays the same: ownership first, tooling second.
This is the work I do with operations teams. If that sounds like your current situation, get in touch.
ITIL trägt den Support über drei Praktiken: Service Desk, Incident Management und Problem Management. Im Betrieb werden diese Praktiken üblicherweise über ein Stufenmodell (L1, L2 & L3) auf die Teams verteilt. Eine saubere Zuordnung der Aufgaben baut ein Team das im Ernstfall einsatzfähig und effizient ist.
Was ITIL unter L1, L2 und L3 versteht
Der Service Desk erfasst die Nachfrage nach Incident Lösung und Service Requests und ist der einzige Kontaktpunkt zwischen Anbieter und Anwendern. Das Incident Management minimiert die negative Auswirkung von Incidents, indem der normale Betrieb so schnell wie möglich wiederhergestellt wird. Das Problem Management senkt Wahrscheinlichkeit und Auswirkung künftiger Incidents, indem es Ursachen identifiziert und Workarounds sowie Known Errors verwaltet.
Eine ungeplante Unterbrechung eines Service oder eine Minderung seiner Qualität.
Eine Ursache oder mögliche Ursache eines oder mehrerer Incidents.
Ein Problem, das analysiert, aber nicht behoben ist.
Eine Lösung, die die Auswirkung reduziert oder beseitigt, solange die vollständige Behebung fehlt.
ITIL definiert Eskalation als das Teilen von Informationen oder das Übertragen der Verantwortung für ein Thema oder Arbeitspaket. Das sind zwei verschiedene Vorgänge in einer Definition. Jemanden informieren ist keine Übergabe der Verantwortung. Wer das nicht trennt, hat am Ende einen Fall, den drei Personen kennen und niemand besitzt.
| Ebene | Praktik in ITIL | Verantwortung | Gibt ab, wenn |
|---|---|---|---|
| L1 | Service Desk | Annahme, Priorisierung, Kommunikation, Lösung nach Runbook | die Meldung ausserhalb des Runbooks liegt |
| L2 | Incident Management, Support-Team | Diagnose, Workaround, Fix, Pflege der Runbooks | die Ursache wiederkehrt oder das Design betrifft |
| L3 | Problem Management, technische Praktiken | Ursachenanalyse, Known Errors, Architektur, Standards | die Behebung eine Änderung über die Change Control braucht |
Ein Support Team konzentriert sich auf die Aufrechterhaltung des Normalbetrieb und löst Anwenderanfragen, Incidents sowie Probleme zu bestimmten Produkten oder Services. Die Zuweisung erfolgt über den Themenbereich des Incidents nicht über die Erfahrung der Person, die gerade greifbar ist.
Ein wichtiger Bestandteil der gerne überlesen wird, ist der Major Incident Prozess. Ein Major Incident ist ein Incident mit erheblicher Geschäftsauswirkung, der eine sofortige koordinierte Lösung verlangt. ITIL nennt dafür Swarming. Viele Beteiligte arbeiten zu Beginn gemeinsam, bis klar ist, wer weitermacht und wer wieder aussteigt.
Ownership
- ZuschnittL1 breit besetzt und nah an den Anwendern. L2 nach Produkt oder Plattform geschnitten, nicht nach Technologie. L3 klein und standardmässig ausserhalb der Bereitschaft. Wenn L3 jede Woche gerufen wird, ist das keine Personalfrage, sondern eine Lücke im Runbook.
- EigentumDas Ticket bleibt bei L1, auch wenn die technische Arbeit wandert. L1 führt die Kommunikation zum Anwender bis zum Abschluss. So bleibt Eskalation eine Übergabe von Arbeit und wird nicht zum Verschwinden des Falls.
- WissenEin Runbook nennt das Dashboard, die Log Zeile, den Befehl und die Prüfung, ob es gewirkt hat. ITIL formuliert es allgemeiner: Wissen ist Information im Kontext dessen, der sie braucht. Ein Handbuch mit 300 Seiten hilft am Service Desk nicht, wenn dort eine Antwort in zwei Minuten fällig ist.
- Eskalation nach der UhrDie Bedingung für die Übergabe steht in der Kategorie und im Timer, nicht im Gefühl der Person am Ticket. Wer erst eskaliert, wenn er aufgibt, eskaliert zu spät.
Wie das in der Praxis aussieht, hängt von der Applikationslandschaft ab und nicht an einer Formel. Als Ausgangspunkt hat sich bewährt: L1 trägt die Fläche und ist durchgehend besetzt. L2 ist pro Plattform mindestens zu zweit, damit Bereitschaft und Tagesgeschäft sich nicht gegenseitig blockieren. L3 besteht aus wenigen Personen, die planbar arbeiten. Die Frage ist nicht, wie viele Ebenen ein Betrieb hat, sondern ob jede Ebene alle notwendige Informationen besitzt, um reibungslos die Problematik angehen zu können.
Das Zusammenspiel braucht drei feste Punkte in der Agenda, sonst zerfällt es in Ernstfall.
Die Übergabe zwischen den Schichten klärt offene Fälle, laufende Workarounds und was in den nächsten Stunden anschlagen kann. Ein wöchentliches Review lässt L1 die Muster melden, die ticketweise unsichtbar bleiben. Aus diesem Verhalten entsteht ein kontinuierlicher Verbesserungsprozess. Das Post Mortem nach jedem grösseren Fall schliesst den Kreis, indem sein Ergebnis im Runbook landet und damit wieder bei L1.
Zwischen L1 und L2 hilft eine benannte Person mehr als eine Gruppe. Ein Mitglied aus L2 ist pro Woche der Ansprechpartner für L1, nimmt Rückfragen und Eskalationen an und hält am Ende der Woche fest, was ins Runbook gehört.
Für den Major Incident braucht es drei benannte Rollen, bevor er eintritt: eine Person, die führt und entscheidet, eine Person, die nach aussen kommuniziert, und die Fachleute, die arbeiten. Wer führt, arbeitet nicht mit. Diese Trennung wirkt im ruhigen Zustand übertrieben und ist im Ernstfall der Unterschied zwischen Koordination und Parallelbetrieb.
Eine Aufgabe gehört zu L1, sobald sie beschrieben, geprüft und begrenzt ist. Beschrieben heisst Runbook. Geprüft heisst einmal gemeinsam durchgeführt. Begrenzt heisst eine Berechtigung, die genau diesen Fall abdeckt und nichts darüber hinaus.
Gemessen wird nicht die Stufe, sondern der Weg. Vier Werte reichen für den Anfang: Anteil der Fälle, die bei L1 enden. Zeit bis zur Eskalation. Anteil wiederkehrender Ursachen. Wiedereröffnungsquote.
ITIL beschreibt den Wassermelonen Effekt. Ein SLA ist aussen grün und innen rot. Die Verfügbarkeit liegt bei 99,6 Prozent, und die fehlenden 0,4 Prozent fallen genau in den Geschäftsvorgang, auf den es ankommt. Wer nur Systemwerte misst, sieht ein Team, das seine Ziele erreicht und Anwender, die etwas anderes berichten.
Die Stufen standen, die Übergaben nicht…
Ausgangslage
In einem Betrieb, den ich übernommen habe, gab es L1, L2 und L3 bereits. Auf dem Papier war das Modell vollständig. Im Alltag lief es zäh.
Die Kommunikation zwischen den Ebenen war mässig. Informationen wurden weitergereicht, aber nicht übergeben. Ein Fall wechselte die Ebene und kam mit denselben Fragen zurück, die vorher schon einmal gestellt worden waren. Ein Teil der Aufgaben war nirgends deklariert. Niemand konnte sagen, ob eine bestimmte Prüfung zu L1 oder zu L2 gehört. Also machte sie mal die eine Seite, mal die andere, mal keine.
Am deutlichsten wurde die Lücke im Krisenfall. Es war nicht festgelegt, wer die Führung übernimmt. Mehrere fähige Leute arbeiteten am selben Fall und Entscheidungen fielen keine, weil jeder annahm, das sei nicht seine Rolle. Die Zeit ging nicht für die Analyse drauf, sondern für die Abstimmung darüber, wer analysiert.
Der Befund
Die Struktur war nicht das Problem. Die Bedingungen waren es.
L1 hatte weder Berechtigungen noch Werkzeuge für Aufgaben, die es fachlich ohne Weiteres hätte übernehmen können. Also wanderte praktisch alles nach oben. Und was oben ankam, kam ohne Kontext an, weil es keine Form gab in der Kontext mitgegeben wird.
Werkzeuge und Berechtigungen
Ich habe die vorhandenen Werkzeuge evaluiert, und zwar nach einem Kriterium: Was verbindet die Ebenen tatsächlich? Gemeinsame Sicht auf die Tickets, Zugriff auf dieselben Dashboards und Logs, nachvollziehbare Aktionen. Nicht ein Werkzeug pro Team, sondern ein Weg durch die Teams.
Danach habe ich die Teams übergreifend befähigt. L1 bekam Zugriff, Rechte und Anleitung für einen festgelegten Satz von Aufgaben, die vorher bei L2 lagen. Für jede dieser Aufgaben galten drei Bedingungen: ein Runbook, das den Fall abdeckt, eine Berechtigung, die genau so weit reicht, und ein definierter Rückweg, wenn es nicht greift. Ohne diese drei Punkte wird eine verschobene Aufgabe zur neuen Fehlerquelle.
Nennt Dashboard, Befehl und die Prüfung, ob es gewirkt hat.
Reicht genau so weit wie die Aufgabe und nicht weiter.
Sagt, wer übernimmt, wenn das Runbook nicht greift.
Der Nebeneffekt war der eigentliche Gewinn. Sobald L1 einen Fall selbst abschliessen konnte, hörte das Weiterreichen als Reflex auf. Wer etwas selbst löst, liest den Fall auch anders.
Post Mortem
Der zweite Schritt war das Post Mortem. Vorher endete ein Fall mit der Wiederherstellung. Danach endete er mit einer Auswertung: was passiert ist, wann es bemerkt wurde, was es verlängert hat, welche Massnahme daraus folgt und wer sie besitzt.
Der Effekt war grösser als erwartet und er lag nicht in der Technik. Das Post Mortem hat Vertrauen aufgebaut nicht nur auf IT Seite, sondern gegenüber dem Business Kunden. Es wurde ohne Schuldzuweisung geführt, so kamen Details auf den Tisch, die vorher unerwähnt geblieben waren. Mit den Details stieg die Qualität der Analyse und mit ihr die Qualität der Massnahmen.
Aus dem Vertrauen entstanden weitere Möglichkeiten. Themen, die vorher niemand angesprochen hätte, wurden zu Einträgen im Verbesserungs Backlog. Der Betrieb bekam damit zum ersten Mal eine Liste, die aus der Realität stammte und nicht aus einer Planungsrunde.
Eingeführt habe ich das Post Mortem nicht als Regel, sondern an Fällen. Die ersten Auswertungen habe ich moderiert und mitgeschrieben. Erst danach wurde es fester Bestandteil. Eine Vorlage vereinfacht den Prozess und stellt sicher, dass die Integrität und Qualität des Post Mortem erhalten bleibt, unabhängig vom Incident.
Mögliche Feldern könnten sein: Zeitleiste, Auswirkung, unmittelbare Ursache, beitragende Faktoren, Massnahmen mit Eigentümer und Termin, und der Eintrag, was davon ins Runbook zurückgeht. Dazu eine Struktur für die Übergabe an L1 und zurück.
Eine Vorlage nimmt niemandem das Denken ab. Sie nimmt die Entscheidung ab, wie das Dokument aussieht und die kostet bei jedem Fall dieselbe Viertelstunde. Nach einigen Fällen waren die Auswertungen vergleichbar. Damit wurden Muster sichtbar, die über mehrere Fälle hinweg liefen und in einem einzelnen Bericht nie aufgefallen wären.
Was nicht auf Anhieb funktioniert hat
Zwei Dinge haben länger gedauert als geplant.
Das erste war die Sorge, etwas kaputtzumachen. Berechtigungen allein ändern kein Verhalten. Die ersten Fälle mussten begleitet werden bis L1 die neue Aufgabe als eigene angenommen hatte. Das zweite war die Versuchung, alles gleichzeitig zu verbessern. Die Verbesserungen nach einem Post Mortem für den L2 in Form von neuen Runbooks war zunächst etwas träge und wurde nicht zeitnah umgesetzt. Geholfen hat, die Restanzen in kleine Pakete und Aufgaben zu verfassen und so inkrementiell, die Verbesserungen einzuführen.
Was sich verändert hat
Der Informationsfluss war der grösste Unterschied. Ein Fall trug seinen Kontext mit, weil es eine Form dafür gab. Die Aufgaben waren deklariert, also entfiel die Abstimmung darüber, wer zuständig ist. Die Eskalationswege wurden kürzer, weil ein Teil der Fälle die Ebene gar nicht mehr wechselte. Und die Eskalationsdauer wurde kürzer, weil eine Übergabe eine Form hatte statt einer Rückfrage.
- Stufen ohne Berechtigungen erzeugen Warteschlangen, keine Entlastung.
- Die Führung im Krisenfall muss benannt sein, bevor die Krise beginnt.
- Das Post Mortem wirkt zuerst auf das Vertrauen und erst danach auf die Technik.
- Eine Vorlage ist kein Formalismus, sondern gesparte Entscheidungszeit pro Fall.
- Was nicht deklariert ist, macht am Ende niemand.
Wo anfangen
Zwei Fragen reichen für den Befund. Erstens: Welche Aufgaben können die verschiedenen Ebenen heute übernehmen, wenn Berechtigung und Runbook vorhanden wären? Zweitens: Wer führt beim nächsten Major Incident und weiss diese Person es?
Wenn eine der beiden Antworten unklar ist, liegt dort die Arbeit. Wenn beide unklar sind, bleibt die Reihenfolge dieselbe: zuerst die Zuständigkeit, dann das Werkzeug.
Genau diese Arbeit mache ich mit Betriebsteams. Wenn das nach der eigenen Lage klingt, nehmen Sie Kontakt auf.
Work together?
I help teams turn ad-hoc operations into resilient, automated systems.