Shift planning
1) Goals and framework
Shift planning ensures continuous platform readiness for incidents and traffic peaks without command burnout. Objectives:- guarantee coverage of SLO-critical processes (deposits, rates/settle, conclusions, KYC/AML);
- Reduce MTTA/MTTR at any time of the day
- comply with labour laws and internal policies (weekends/breaks/nights).
2) Roles and support levels
L1 (NOC/Operations): primary triage, running runbook, escalation.
L2 (Domain On-Call): Payments, Games/Core, Data/Infra - deep diagnostics and fixes.
L3 (SRE/Platform/Dev): configuration changes, patches, emergency releases.
IC/CL (Incident Commander/Comms Lead): Manages incident and communications.
Duty Manager (by shift): confirms staffing, risks, freeze-windows.
3) Shift and rotation models
3. 1 24 × 7 coverage
8 × 3 (three eight-hour): Day (08: 00-16: 00), Swing (16: 00-00: 00), Night (00: 00-08: 00). Easy to manage, requires more employees.
12 × 2 (two twelve-hour): Day (08: 00-20: 00), Night (20: 00-08: 00). Fewer handovers, higher risk of fatigue - use with restrictions.
Follow-the-sun: EU → AMER → APAC, minimal night; requires a distributed command.
3. 2 Rotation templates (example)
Panama/Pitman (2-2, 3-2, 2-3): alternating 12-hour with long weekends (requires strict fatigue control).
4-on/4-off (12h): 4 days for 12 hours, 4 days off; suitable for L1 with good automation.
5 × 8 (classic): for L2/L3 + on-call nights/weekends.
3. 3 On-call duty
Primary/Secondary: Each domain has a primary and secondary on-call.
Shadow-on-call: Training new engineers under Primary's wing.
4) Calculation of state and shrinkage (shrinkage)
4. 1 Basic FTE formula
Required FTEs for the role:
FTE = (coverage hours per week/productive FTE hours per week) × (1 + shrinkage)
Where shrinkage includes vacation/sick leave/training/1: 1/retro/admin time (usually 20-35%).
Example (L1 24 × 7, 8 hours):- 168 hours coverage/35 productive hours/week ≈ 4. 8
- With shrinkage 30% → 4. 8 × 1. 3 ≈ 6. 2 FTE (round to 7 for stability).
4. 2 Peak plan
Add peak factor for events (top matches, tournaments): + 10-25% FTE in calendar windows. Use historical alert/traffic plots.
5) SLO coverage and risk windows
Build a matrix of critical hours (local prime time GEO/payment methods).
Set the minimum composition per slot (for example, Night: L1 × 2, L2-Payments × 1 on-call, L2-Games × 1 on-call, IC by rotation).
For releases/migrations, assign a release guard (option L2/SRE) for the period and + 60 minutes of post-monitoring.
6) Handovers (shift transfer)
Structure 10-15 minutes:1. SLO/alert and open incident status.
2. Planned work in the window + risks.
3. Blockers/waits (PSP/KYC providers, regulation).
4. Agreed comm plans (status page, partners).
5. Check shift composition and on-call contacts.
The handover checklist must be in the wiki/bot; protocol - in var-room or replaceable channel.
7) Calendar and replacements
Planning horizon: 8-12 weeks; replacement - no later than 2 weeks (except for force majeure).
Freeze periods: major events/holidays - vacation ban for key roles (compensated later).
Buddy-rule: replacement only by an engineer of the same domain/qualification or with additional shadow.
8) Platform integrations
Alerting: routing by active shift and domains (P1→pager+var room).
Incident bot: teams '/rota ', '/whoisoncall', auto-mentions IC/CL.
Releases: CAB synchronized with shift schedule; off-coverage releases - prohibited.
Status page: CL from active shift confirmed in bot.
9) Team sustainability and health
Labor standards: night/weekend - surcharges and time off; maximum night in a row (for example, ≤3).
Breaks: every 2-3 hours a short break; at 12h - mandatory two long.
Fatigue-watch: limit of hours/week, the rule "no more than N P1 per shift per one."
Psychological support: debrief after heavy P1, comp-time.
10) Multi-region and follow-the-sun
Divide brands/tenants by region (EU/LATAM/APAC) with local L1 and domain L2.
Regional ICs escalate to a global IC in cross-regional incidents.
Daily cross-region handover (15 min) with transfer of risks and works.
11) Politicians and RACI
Policy "Shift & On-Call": Who steps in, how and when; replacements; tardiness; reserve.
SoD/accesses: IC/CL/financial transactions - separate roles; JIT boosts via bot only.
RACI: each shift has IC (A), Duty Manager (A/R), L1/L2 (R), Compliance/Sec (C), management (I).
12) Tools and data
Single calendar (with attributes: domain, region, contact, reserve).
Shift load dashboards: alerts/hour, incidents/type, releases/windows.
Justice Reports (fair-share): Night/weekend coverage by people.
Integration with HR/PTO so that shrinkage is counted automatically.
13) Quality Metrics (KPI/KRI)
Coverage Rate:% hours with full cast.
MTTA/MTTR by slot: day/evening/night.
Handover Defects: number of missing checklist items.
Pager Fatigue: alert/person/week; night calls (target - ↓).
Replacement SLA:% of closed shifts by replacement ≤ 48 hours before the start.
Training Coverage: share of shifts with shadow slot for new operators.
Fair-Share Index: Evenness of nights/days off by employee.
14) Implementation Roadmap (4-8 weeks)
Ned. 1-2: collect the history of alerts/peaks, determine SLO-critical windows; choose a model (8 × 3 or follow-the-sun), calculate FTE and shrinkage.
Ned. 3-4: publish Policy "Shift & On-Call," enable handover-checklist, run general calendar and bot commands '/whoisoncall ', '/handover'.
Ned. 5-6: debug escalations and replacements, add fair-share reports, synchronize CAB/releases with shifts, enter freeze windows.
Ned. 7-8: retro in terms of burnout and quality of handovers, correction of the composition of night ones, launch of a shadow program and IC/CL certification.
15) Patterns and artifacts
Shift Rota (example for EU, Kyiv time):Handover Checklist (10 points): SLO, incidents, works, risks, releases, providers, accesses, status pages, open tickets, staffing.
Replacement SOP: how to issue a replacement, admission criteria, reserve contact.
Freeze Calendar: events/holidays and bans on RTO/releases.
16) Antipatterns
One on-call for all without domain separation.
12-hour nights in a row without restrictions and breaks.
Releases in hours without IC/CL/CL availability.
Handover "on spoken word," without checklist and notes.
Ignore shrinkage → a chronic under-set.
Zero shadow program → fragility and SPOF in humans.
Total
Shift planning is an engineering task: calculation of FTE and shrinkage, honest rotation and health protection, strict handovers and integration with alerting/ChatOps/CAB. This framework provides predictable 24 × 7 SLO coverage, fast incident responses, and business resilience without team burnout.