Principal Hardware Diagnostics Engineer

GraphcoreMilpitas, CA
1d

About The Position

We are seeking an experienced Principal Hardware Diagnostics Engineer to design and develop diagnostics software used to monitor hardware health and diagnose system-level issues across Graphcore’s AI infrastructure platforms. This role focuses on building diagnostics agents, tools, and analytics frameworks that enable engineers and automation systems to identify, isolate, and resolve hardware issues across blade-level servers and rack-scale clusters.

Requirements

  • Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, or related discipline.
  • Strong software engineering experience in Python, C++, or C#.
  • Experience developing diagnostics or monitoring systems for hardware platforms.
  • Experience working with distributed systems or cloud infrastructure.
  • Strong knowledge of Linux environments and system-level diagnostics tools.
  • Experience collaborating with CM/ODM partners on manufacturing diagnostics and fault isolation.
  • Strong analytical and debugging skills.
  • Excellent communication and collaboration abilities.

Nice To Haves

  • Experience working with AI hardware platforms or accelerator-based computing systems.
  • Familiarity with hyperscale data center infrastructure.
  • Experience building cluster-level monitoring or diagnostics systems.
  • Experience interacting with internal or external customers during diagnostics solution development.

Responsibilities

  • Design and develop automated hardware diagnostics solutions for blade-level servers and rack-scale AI systems.
  • Architect and implement diagnostic agents, monitoring tools, and analytics frameworks to track hardware telemetry.
  • Collaborate with hardware teams to integrate low-level diagnostic modules into monitoring systems.
  • Develop diagnostics tools capable of detecting hardware health conditions and isolating failures.
  • Create diagnostic modules used for internal validation and production data center operations.
  • Provide detailed hardware fault information to system engineers to accelerate troubleshooting.
  • Define remediation workflows and insights for hardware fault scenarios across nodes and clusters.
  • Collaborate with firmware, networking, and cloud platform teams to integrate diagnostics across the system stack.
© 2024 Teal Labs, Inc
Privacy PolicyTerms of Service