Job Title: Fleet Debugging & Incident Management Engineer
Key areas to look for include:
Server platform validation/SDET/FW development etc (BIOS, BMC, DPU,GPU, storage, networking, server bring-up)
Fleet diagnostics, root cause analysis, and issue triage
Telemetry, monitoring, dashboarding, and log analysis
Python/PowerShell automation and operational tooling
Azure DevOps, incident management, and release validation
Experience supporting server cloud infrastructure, datacenter, or large-scale deployment environments
Job Summary
We are seeking a proactive engineer to support large-scale fleet operations, incident management, and platform debugging activities. The role requires ownership of production issues, cross-functional coordination, root cause investigations, and executive-level reporting to ensure fleet health and operational readiness.
Key Responsibilities
· Own ICM triage, incident tracking, escalation, and closure.
· Perform fleet debugging across hardware, firmware, BMC, BIOS, and platform software layers.
· Analyze logs, telemetry, crash dumps, and diagnostics to identify failure patterns and root causes.
· Coordinate with platform, validation, firmware, release, and deployment teams to drive issue resolution.
· Lead operational reviews and provide status, risk, and incident trend reporting to stakeholders.
· Work closely with offshore engineering teams, providing technical guidance and execution oversight.
· Support release readiness assessments and fleet health monitoring activities.
· Identify opportunities for automation and operational process improvements.
Required Skills
· Robust communication, stakeholder management, and problem-solving skills.
· Strong experience in hardware and firmware debugging.
· Hands-on knowledge of server platforms, BMC, BIOS, Linux, and low-level system architecture.
· Experience with incident management, production support, or fleet operations.
· Ability to analyze logs, telemetry , and system diagnostics to troubleshoot complex issues.
· Proficiency in Python, Shell, or PowerShell scripting.
Preferred Skills
· Experience with DPU, GPU, accelerator, or cloud-scale server platforms.
· Exposure to Azure infrastructure, data center operations , or reliability engineering.
· Familiarity with ICM, monitoring systems.
· Experience leading investigations across geographically distributed teams.
Key Focus Areas: Fleet Debugging, ICM Management, Root Cause Analysis, Platform Reliability, Operational Reporting, Cross-Team Coordination, and Incident Resolution.
Thanks & Regards,
Trayambkeshwer Dwivedi (Trayam), Sr. Technical Recruiter
Raas infotek corporation
262 Chapman road, Suite 105A, Newark, DE-19702
Direct number: (phone hidden) | 132
Text Now: (424) 222 7980
Email :
[email protected]
📌 Fleet Debugging & Incident Management Engineer (Canada)
🏢 Raas Infotek
📍 Canada