07 Aug
|
Creative Solutions Services
|
Mississauga
07 Aug
Creative Solutions Services
Mississauga
Global Financial Firm located in MISSISSAUGA, ON has an immediate contract opportunity for an experienced Production Support AnalystThis role is currently on a Hybrid Schedule. You will need to have reliable internet, computer and android or iphone for remote access into the client systems during remote work. We will be expected in the office weekly 3 days depending on the team requirement.Video/ f2f interviews are required prior to all offers.Pay rate range: $ 70.00 - $ 75.00 Negotiable based upon years of experienceJob Description:Role OverviewThis is a senior technology operations leadership role responsible for the stability, resilience, data integrity, and continuous improvement of the Enterprise Risk Management (ERM) platform — a high-criticality, large-scale application comprising 17+ sub-pillars and a growing portfolio of agentic AI modules. The role sits at the intersection of production support, platform engineering, DevOps, and data platform operations, requiring deep technical capability across application support, event-driven data pipelines, and enterprise data architecture combined with strong operational discipline and cross-team coordination.The successful candidate will lead an L3 engineering function, drive DevOps maturity, own data platform operations (including ERDL and the ERM Data Lake), and serve as the technical bridge between development teams, data engineering, infrastructure, and business stakeholders. This is not a passive support role — it is an active platform ownership position with accountability for production health, data quality, release governance, incident resolution, and engineering excellence.ResponsibilitiesProduction Support & Incident ManagementLead L3 triage and resolution of complex production incidents across a 17+ pillar enterprise risk platform, including data ingestion failures, workflow disruptions, Kafka messaging issues, and infrastructure events.Own the Problem Record (PRB) lifecycle — from triage and root cause analysis to fix coordination and post-incident documentation — in alignment with ServiceNow ITSM processes.Drive the SWAT process: daily review of open problem tickets with escalation potential, ensuring senior stakeholder awareness and timely resolution.Serve as the primary escalation point for L3 engineers, coordinating with development, middleware, DBA, and Tenant Ops teams as needed.Lead Major Incident Management (MIM) for high-impact production events including data pipeline failures, SLA misses, and reconciliation discrepancies.Data Platform Operations & ArchitectureOwn L3 production support for the Enterprise Risk Data Layer (ERDL) — the central data platform that aggregates risk data from 17+ upstream Pillar systems and Federated Limits Units (FLUs).Triage and resolve data pipeline failures across Kafka, Oracle, and the ERM Data Lake — including ingestion errors, materialized view refresh failures, batch job timeouts, and data reconciliation discrepancies.Perform and govern daily data quality checks: KRI count reconciliation (ERDL vs. KRI API vs. KRI/RMD Dashboard), RA count reconciliation, PRIVATE_IND NULL checks, full and incremental refresh status monitoring, and Kafka ingestion lag monitoring.Maintain deep working knowledge of the ERDL reporting layer hierarchy and its authoritative use cases:RMD (Risk Management Dashboard): Authoritative for business user consumption and management reportingERDL Recon View: Authoritative for reconciliation and data quality checksERDL Tableau Extract: Authoritative for historical as-of views for reconciliation purposesIdentify, classify, and escalate SLA Miss PRBs — distinguishing ERM application defects from upstream FLU compliance failures, and maintaining a running log of SLA miss frequency by FLU for governance escalation.Support and govern the ERM Data Lake — understanding data flows, retention policies, and downstream consumer dependencies.Data Modeling & Data LineageMaintain and apply working knowledge of the ERM canonical data model — understanding key entities (overlays, limits, thresholds, KRIs, risk assessments), their relationships, and how they flow through the platform.Understand and document end-to-end data lineage for critical data flows: from upstream FLU systems through Kafka, into ERDL, through to downstream consumers (RMD, Tableau, Superset,
Recon).Apply data contract principles to triage integration failures — identifying where producer/consumer misalignments (missing fields, incorrect data types, null values, schema mismatches) are causing production defects.Contribute to data architecture governance: support the validation of AI-generated data contract analyses, participate in Data Contract Review meetings, and help translate findings into Jira backlog items.Classify production problems accurately as either Implementation Bugs (coding/configuration errors) or Policy/Process Misalignments (failures to correctly execute business rules), articulating the distinction clearly in PRB documentation.Support the ERDL Refactoring initiative and related data architecture modernization efforts, including Redis infrastructure upgrades and Prod Parallel Deployment stabilization.Platform & DevOps EngineeringPerform and coordinate DevOps functions across lower environments (DEV, SIT, UAT), including environment management, deployment sequencing, and release validation.Lead or coordinate weekly production releases — owning the release bridge, executing health validation in Harness and OpenShift, and ensuring complete post-deployment documentation.Drive CI/CD pipeline health across Client Project and GitHub — ensuring build integrity, scan compliance (Snyk, SonarQube, Checkmarx), and deployment readiness across all Pillars.Manage Continuous Vulnerability Management (CVM) across the platform, coordinating remediation plans with development teams by Pillar.Lead cross-pillar DevOps initiatives: Angular/React upgrades, Python version decommissions, Tomcat upgrades, Hashicorp Vault onboarding, SSL certificate lifecycle management, and GitHub migration.Monitoring, Observability & AutomationOwn and evolve the daily operational monitoring framework — Kafka ingestion health, data reconciliation, refresh status, and data quality checks.Drive automation of manual monitoring activities, including automated ServiceNow incident creation for detected anomalies.Leverage AppDynamics, Kafka dashboards (Tableau/Superset), and OpenShift tooling for real-time platform health visibility.Identify and close monitoring gaps — including proactive FID/AD group membership monitoring to prevent silent infrastructure failures.Governance, Standards & Engineering ExcellenceEnforce L3 operational standards: PRB description quality, PTASK lifecycle compliance, and Manual Touch Point (MTP) process adherence.Champion Developer Manifesto compliance: README standards, GitCode ownership, branch hygiene, stale repository cleanup, and CI/CD health metrics.Ensure compliance with technology risk, security, IS assessment, and regulatory standards (GIAM, EERS, CVM, CAMP, DPS data protection standards).Apply the AI-Assisted Analysis and Governance framework to accelerate root cause identification and translate findings into auditable Jira backlogs.Strategic & Stakeholder LeadershipServe as the operational owner for new application onboarding (e.G., OMAI Overlay, Tapas, Shock Generation AI) — defining support models, escalation matrices, and runbooks.Represent BAU production health and data platform status at weekly governance calls, cross-pillar bi-weekly calls, and SWAT touchpoints.Partner with Architecture, Product Owners, Development Leads, Data Engineers, and Infrastructure teams to coordinate cross-pillar initiatives.Mentor L3 engineers; foster a culture of operational excellence, data quality ownership, and structured knowledge sharing.Qualifications8+ years of technology experience, with a strong background in production engineering, L3 support, data platform operations, or platform DevOps in a large-scale enterprise environment.Proven experience in a senior technical operations or engineering lead role with accountability for production stability, data pipeline health, and team coordination.Demonstrated ability to triage and resolve complex,
multi-system production issues across distributed microservices and data pipeline architectures.Solid hands-on experience with data platform operations — including event-driven pipelines, data reconciliation processes, and multi-layer reporting architectures.Working knowledge of data modeling concepts — entity relationships, canonical data models, schema evolution, and data contract principles.Experience with data lineage analysis — tracing data flows from source systems through transformation layers to downstream consumers and identifying break points.Strong experience with ITSM processes (ServiceNow — incident, problem, change, MTP/PRJ modules) in a formal IT governance environment.Experience coordinating production releases — runbook execution, health validation, and stakeholder communication.Excellent communication, escalation management, and stakeholder engagement skills — comfortable representing technical and data quality status to senior leadership.Experience in financial services or regulated technology environments strongly preferred.Technical SkillsData Platform & ArchitectureKafka / JMS — producer/consumer health monitoring, topic-level triage, schema validation, consumer lag analysisOracle — read-level query capability; working knowledge of materialized views, batch jobs, schema structures, and data reconciliation viewsData lineage tooling — ability to trace and document data flows across multi-system architecturesData contract principles — producer/consumer responsibilities, schema validation, null-safety, field-level contract analysisTableau / Superset — operational dashboard monitoring and data layer reconciliationElastic Search — basic operational awareness for search/index layer triageMongoDB / Couchbase — operational awareness for NoSQL data stores in use across PillarsData Lake concepts — retention policies, data classification, downstream consumer patternsPlatform & InfrastructureOpenShift / Kubernetes — pod management, health checks, container operationsHarness — deployment pipeline management and release validationClient Project (LSE / Classic) — CI/CD platform managementGitHub / Bitbucket — repository governance, branch management, pipeline configurationAppDynamics — application performance monitoring and alertingApplication & IntegrationJava / Spring Boot — sufficient depth to triage application-layer issues, interpret stack traces, and understand data contract failuresREST APIs — API failure interpretation, connectivity validation, integration troubleshootingAngular / React — basic familiarity for front-end issue triageObservability & ToolingAppDynamics, Splunk, ELK / Kibana — log analysis and alertingSonarQube, Snyk, Checkmarx — compliance gate interpretationServiceNow — incident, problem, change, and PRJ module managementSecurity & ComplianceCyberArk, CISAR — FID and privileged access managementEEMS / EERS — entitlement management and access review processesCVM / CAMP — vulnerability management and Pillar-level remediation coordinationHashicorp Vault — secrets management operationsSSL / TLS certificate lifecycle managementPreferredExperience with ERDL or equivalent enterprise risk data layer platforms — aggregating data from multiple upstream systems into a central risk reporting layer.Familiarity with Kafka schema governance and event-driven integration patterns between enterprise risk systems (limits, thresholds, KRIs, model risk).Experience supporting or onboarding agentic AI or GenAI-integrated applications (e.G., Generative AI overlays, LLM-backed workflows, MCP-based architectures).Experience implementing automated data quality monitoring and self-healing alerting frameworks.Knowledge of FAST automation framework, contract testing, or behavior-driven development (Gherkin/Cucumber/Selenium).Familiarity with cloud modernization (Cloud @ Client / Type A migration) and container-native platform evolution.Exposure to enterprise risk management concepts — 1LOD/2LOD governance, stress testing (CCAR/QMMF), model risk management, or limits and thresholds frameworks.Experience with Redis infrastructure for caching and entitlement stability.EducationBachelor's degree or equivalent experience in Computer Science, Engineering, Information Technology, Data Engineering, or a related field. #J-18808-Ljbffr
📌 Production Support Analyst (Mississauga)
🏢 Creative Solutions Services
📍 Mississauga