Please be advised that our Careers site will be unavailable from November 28 at 12am ET to November 29 12am ET for scheduled system maintenance.

Title:  Systems Reliability Engineering Senior Manager

 

 

 

 

Requisition ID: 272518

Thanks for your interest in ScotiaTech, Scotiabank's new and innovative Technology hub in Bogota.

Join a purpose driven winning team that promotes creativity and innovation in a fast-paced environment, where we’re always committed to results, in an inclusive, diverse, and high-performing culture.

 

Purpose

Contributes to the success of International Banking by leading the operation of the Monitoring HUB, ensuring that the Level 5 Analyst team executes monitoring, triage, alert handling, handovers, and escalations in accordance with established procedures. The role is responsible for operational coordination, administrative management, team performance, SLA/OLA tracking, financial control, executive reporting, and the continuous improvement of the monitoring model. All activities are carried out in compliance with applicable regulations, internal policies, control culture, operational risk management, security, compliance requirements, and standards of conduct.

 

Accountabilities 

Lead the daily operation of the SRE Analyst Level 5 team, ensuring coverage, prioritization, operational discipline, incident follow-up, and adherence to established playbooks.

Manage shift scheduling, night rotation, extended coverage, backfills, vacations, absences, handovers, and the operational continuity of the monitoring team.

Define and maintain the team’s operational calendar, ensuring workload balance, assignment traceability, and alignment with business needs.

Oversee the monitoring of International Banking applications and services, ensuring that monitoring screens, alerts, dashboards, GEMS, Dynatrace, Grail, logs, and ITSM tools are reviewed according to the operating model.

Ensure alerts, service degradations, and incidents receive timely initial triage, sufficient evidence collection, preliminary classification, appropriate escalation, and follow-up through operational closure or handover.

Coordinate operational priorities with application, infrastructure, incident management, business operations, and SRE Central teams when support, decisions, or remediation fall outside the Level 5 scope.

Escalate opportunities and priorities to SRE Central related to observability, alerting, dashboards, thresholds, operational noise reduction, automation, remediation, playbooks, and regional standardization.

Propose improvements to alerting, dashboards, operational views, event correlation, service health monitoring, detection metrics, and procedures to reduce operational noise and improve MTTD, MTTR, and triage quality.

Govern the quality of shift handovers, operational logs, incident evidence, checklist compliance, and updates to operational documentation.

Manage team performance metrics, including shift coverage, handover compliance, alerts reviewed, incidents triaged, timely escalations, operational backlog, SLA/OLA adherence, documentation quality, and procedural compliance.

Prepare executive and operational reports on team performance, KPIs, capacity, allocation, costs, risks, alerting trends, recurring issues, and improvement initiatives.

 Manage team costs, capacity planning, resource allocation, headcount requirements, productivity, operational budget, and forecasting required to sustain Monitoring HUB coverage.

Oversee administrative activities for the team, including onboarding, training, required access provisioning, knowledge development plans, compliance tracking, operational meetings, process feedback, and coordination of onsite activities.
Promote a strong control culture, compliance with procedures, operational risk management, information security, and responsible use of tools and access privileges.
Represent the monitoring team in operational forums, service reviews, incident committees, HUB governance meetings, and stakeholder follow-up sessions within International Banking.
Foster an inclusive, collaborative, and high-performing work environment focused on service excellence, continuous learning, operational reliability, and continuous improvement.

 

Education / Experience / Other Information 

•    Bachelor's degree in Systems Engineering, Computer Science, Telecommunications, Technology Management, Industrial Engineering, or a related field. Equivalent experience in critical IT operations may be considered.
•    Strong experience leading teams in operations, monitoring, NOC, SRE, incident management, application support, production support, or high-availability technology services.
•    Experience managing shift schedules, coverage models, operational continuity, staffing, capacity planning, resource allocation, cost control, and on-site or hybrid teams.
•    Practical knowledge of observability and monitoring tools, including Dynatrace, dashboards, alerts, logs, metrics, traces, GEMS, Grail, ITSM, and comparable platforms.
•    Ability to interpret operational and executive metrics, identify trends, prioritize improvements, and present information clearly to both technical and non-technical stakeholders.
•    Knowledge of ITIL or equivalent processes, including incident management, problem management, escalation management, service reviews, operational governance, and continuous improvement.
•    Knowledge of Site Reliability Engineering (SRE) management practices, including SLIs, SLOs, error budgets, toil reduction, reliability reporting, observability standards, and operational readiness.
•    Experience proposing improvements to alerting, dashboards, playbooks, runbooks, recurring event management, thresholds, and triage procedures.
•    Strong leadership, executive communication, organizational, negotiation, prioritization, conflict resolution, decision-making, and follow-up skills.
•    English proficiency is preferred for communication with regional or global teams. Candidates without English proficiency must demonstrate the ability to perform the role in Spanish and leverage established channels for interactions requiring English.
•    Availability to work on-site in Bogotá and to support critical events or coordination needs outside standard business hours when required by the operating model.

 

Working Conditions

On-site role in Bogotá, Colombia, leading the Level 5 Monitoring Team from the office.
Standard schedule is Monday through Friday, from 8:00 a.m. to 5:30 p.m., with flexibility required to support operational reviews, major incidents, shift changes, and critical escalations as needed.

 

#LI-HYBRID


Location(s):  Colombia : Bogota : Bogota 

ScotiaTech is a business unit within ScotiaGBS, a Scotiabank Group company located in Bogota, Colombia. The ScotiaTech hub was created to support different technology systems and processes of the Bank. We offer an inclusive, positive work environment, and competitive benefits.

At ScotiaTech, we value the unique skills and experiences each individual brings and are committed to creating and maintaining an inclusive and accessible environment for everyone. Candidates must apply directly online to be considered for this role. We thank all applicants for their interest in a career at ScotiaTech; however, only those candidates who are selected for an interview will be contacted.

Note: All postings in me@Scotiabank will remain live for a minimum of 5 days.


Job Segment: Telecom, Telecommunications, Risk Management, Systems Engineer, Industrial Engineer, Technology, Finance, Engineering