
Cluster Datacenter Manager
United States | Alabama at a glance
- Rent
- #40 of 511-bedroom-31% vs US avg
- Weather
- #10 of 51184 mild daysstatewide median
- Income tax
- #23 of 515% top rateAlabama
What you need
- 10+ yrs data center/compute infrastructure experience
- 5+ yrs leading technical teams
- Multi-site operations management experience
- Hands-on server hardware, rack, diagnostics expertise
- Data center network topology, fiber troubleshooting skills
What you'll do
- Own compute cluster operational health across sites
- Lead hardware incident triage, break/fix, FRU replacement
- Lead fiber/copper physical-layer troubleshooting
- Manage asset lifecycle, secure media handling
- Lead technicians, engineers across multiple facilities
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
The Role
Cerebras is seeking a hands-on Datacenter Cluster Manager to lead compute infrastructure operations and operational reliability across multiple data center sites. This leader manages technicians, engineers and contractors responsible for the health, availability, physical maintenance and lifecycle of high-performance compute and network infrastructure. The role requires strong technical judgment in server systems, networking, fiber connectivity and secure asset handling, alongside a working understanding of power, cooling and colocation dependencies. The Cluster Manager leads incident response, drives operational discipline, and ensures new deployments transition safely into sustained production support.
Responsibilities
Compute Systems Operations and Reliability
Own day-to-day operational health and service readiness of compute clusters, server racks, supporting network infrastructure and associated physical systems across assigned sites.
Lead hardware incident triage, fault isolation, break/fix, field-replaceable unit (FRU) replacement, diagnostic validation and escalation to systems engineering or vendors.
Use monitoring, telemetry, logs and alerting to recognize degraded systems, prioritize impact, coordinate recovery and verify restoration of service.
Oversee preventive maintenance, firmware or hardware change execution where authorized, spare parts readiness, and repeat-failure analysis.
Partner with cluster ops , network and reliability engineering teams on root cause analysis, corrective actions and recurring fleet health issues.
Network Infrastructure and Fiber Troubleshooting
Apply practical understanding of data center network architecture, rack-level connectivity, switching, management networks and network redundancy to guide onsite diagnosis and escalation.
Lead physical-layer troubleshooting of fiber and copper links, including patching, optics and transceivers, polarity, cleanliness, labeling and continuity testing using appropriate tools.
Ensure structured cabling, fiber routing, documentation and change controls are maintained to engineering standards.
Coordinate with network engineering to isolate physical connectivity faults versus configuration or software issues; execute approved remediation without assuming ownership of network design.
Secure Media, Assets and Hardware Lifecycle
Own physical asset accountability from inbound receiving, inspection, staging and inventory through deployment, repair, return and disposition.
Enforce secure handling of data-bearing devices, including authorization, chain of custody, access controls, approved sanitization or destruction workflows and documented transfer.
Ensure adherence to customer-specific controls, including two-person authorization and clean-in/clean-out procedures where applicable.
Maintain accurate rack elevations, asset records, serial numbers, spares, repair histories and inventory reconciliation.
People Leadership and Cluster Execution
Lead, coach and evaluate technicians, engineers and contractor resources across multiple facilities; establish clear ownership, training and performance expectations.
Plan staffing, shift coverage, on-call rotations and escalation readiness based on cluster demand, production service commitments and deployment schedules.
Prioritize work across locations, protect production operations during concurrent builds and maintain consistent SOPs, MOPs, ticketing and change management.
Own incident command, stakeholder communications, post-incident reviews and measurable corrective action tracking.
Deployment, Expansion and Production Handover
Oversee operational readiness for new compute capacity, including receiving, rack integration, network and fiber validation, power-on checks, system health verification and production acceptance.
Coordinate installation, commissioning, vendor work and handover with deployment, infrastructure, engineering and program teams.
Confirm documentation, tooling, spares, staffing, access and escalation paths are ready before accepting operational ownership.
Critical Facilities, Vendors and Operational Controls
Understand dependencies between compute availability and electrical distribution, UPS, generators, PDUs/RPPs, HVAC, liquid cooling, CDUs and environmental monitoring.
Coordinate with colocation providers and facilities specialists on maintenance risk, alarms, degraded cooling or power conditions and incident restoration.
Manage vendor performance, site safety, access controls, maintenance windows, service commitments, operating expenses and contractor utilization.
Required Qualifications
10+ years of experience in data center, compute infrastructure, server operations or related technical operations, including 5+ years leading technical teams.
Demonstrated experience managing multi-site operations or complex data center environments with 24/7 production support expectations.
Strong hands-on foundation in server hardware, rack infrastructure, hardware diagnostics, break/fix and systems health monitoring.
Working understanding of data center network topology and physical connectivity, including Ethernet, fiber optics, transceivers and structured cabling troubleshooting.
Experience with asset management, data-bearing media security, chain of custody and controlled equipment movement.
Experience leading incidents, change control, technical escalations, vendor coordination and deployment-to-operations handovers.
Ability to interpret facility power and cooling dependencies and engage subject matter experts to manage compute service risk.
Clear communication, sound technical judgment and experience developing technicians and engineers.
Preferred Qualifications
Experience supporting large compute clusters or high-density server environments.
Exposure to direct liquid cooling, CDU operations, thermal telemetry and the relationship between cooling performance and system health.
Familiarity with Linux-based troubleshooting, BMC/IPMI or equivalent out-of-band management, server logs and hardware telemetry.
Experience using fiber inspection and testing tools, network test equipment and disciplined physical-layer troubleshooting procedures.
Experience with Grafana, Prometheus, PagerDuty, Jira, DCIM/BMS or equivalent monitoring and workflow tools.
Knowledge of customer security requirements, colocation operating models, SLAs and relevant compliance controls.
What Success Looks Like
Healthy, available compute and network infrastructure with fast detection, structured triage and durable resolution of recurring failures.
Consistent physical-layer troubleshooting, secure media handling and accurate hardware and asset records across the cluster.
Skilled onsite teams with reliable coverage, clear escalation ownership and consistent operational execution.
New compute capacity enters production with verified readiness, documented ownership and minimal impact on live services.
Effective coordination with facilities partners on power and cooling risks without losing focus on compute operations.
Working Expectations
Regular onsite presence and travel between assigned cluster facilities, with additional travel for deployments or operational priorities.
Availability to lead critical incidents, scheduled maintenance and deployment activities outside standard hours when needed.
Specific location, reporting line, travel expectations and compensation to be confirmed for each requisition.
Why Join Cerebras
People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:
Build a breakthrough AI platform beyond the constraints of the GPU.
Publish and open source their cutting-edge AI research.
Work on one of the fastest AI supercomputers in the world.
Enjoy job stability with startup vitality.
Our simple, non-corporate work culture that respects individual beliefs.
Find out more about what it's like to work at Cerebras here!
Apply today and become part of the forefront of groundbreaking advancements in AI!
Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.
This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.
Optimize your resume for this job
Get a match score and the keywords you're missing
About Cerebras
Cerebras Systems builds the world's fastest AI computers. Their Wafer Scale Engine is the largest chip ever built, powering a new class of AI supercomputers that accelerate training and inference by orders of magnitude.
Similar Operations roles


