About the Role
Nscale is seeking an Engineer with strong people, leadership, and technical skills to ensure the efficiency, reliability, and scalability of data center infrastructure. This role requires comfort in problem-solving complex, ambiguous topics in a results-driven environment, influencing without authority, and building senior stakeholder relationships. The ideal candidate quickly grasps technical concepts, possesses strong analytical skills, is organized, diligent, a self-starter, curious, and a quick learner.
Responsibilities
- Join the Support duty rotation and handle day-to-day tickets and alerts, escalating early and appropriately.
- Collaborate with Engineering for incident or change management.
- Accurately record, update, manage, and resolve tickets using the ticketing system, keeping all parties informed.
- Follow established runbooks to resolve common issues and propose improvements.
- Keep tickets updated with clear notes, next steps, and customer communications.
- Learn Platform fundamentals to assist customers and ask for support when needed.
- Participate in monitoring, troubleshooting, and triage, capturing logs and facts for efficient handover.
- Deliver assigned tasks and project work to agreed quality and timelines, flagging blockers early.
- Share knowledge by documenting validated steps and contributing to training materials.
- Shadow seniors during complex work to build capability.
- Participate in incident reviews and track preventative follow-ups.
- Identify areas for automation to optimize processes.
- Continuously learn and upskill.
- Collaborate with cross-functional teams for service improvements.
- Act as the escalation point for onsite operations staff.
- Participate in on-call or out-of-hours work after onboarding.
- Travel to Nscale or Customer locations for deployments, troubleshooting, and operational tasks.
- Attend supplier-related training courses.
Requirements
- 2–4 years of experience in a support, operations, or infrastructure engineering role, ideally within a cloud, data center, or managed services environment.
- Growth mindset: curious, dependable, and collaborative.
- Awareness of servers, networks, storage, and virtualization concepts.
- Comfortable with Linux CLI, systemd services, filesystems, permissions, and basic networking tools.
- Ability to troubleshoot common Linux issues and know when to escalate.
- Solid grasp of IP addressing, subnets, VLANs, routing at a high level, DNS, and firewalls.
- Understanding of core Kubernetes concepts (nodes, pods, services, logs).
- Ability to perform basic Kubernetes troubleshooting and follow runbooks.
- Familiarity with basic GPU diagnostics (e.g., nvidia-smi).
- Ability to use dashboards and alerts to identify symptoms, gather evidence, and follow runbooks.
- Comfortable proposing simple alert or dashboard tweaks with review.
- Comfortable reading and writing simple Bash or Python snippets.
- Experience using Git for version control.
- Familiarity with common hypervisor or cloud troubleshooting flows.
Skills
- People skills
- Leadership skills
- Technical skills
- Problem solving
- Decision making
- Influencing without authority
- Relationship building
- Analytical skills
- Organization
- Diligence
- Self-starter
- Curiosity
- Quick learning
- Linux fundamentals
- Networking basics
- Kubernetes exposure
- GPU awareness
- Observability foundations
- Scripting and automation basics (Bash, Python)
- Git
- Cloud and virtualization basics
- Kubernetes administration (Nice to Have)
- Operators (Nice to Have)
- Storage or networking add-ons (Nice to Have)
- Deeper GPU/HPC concepts (Nice to Have)
- RDMA/InfiniBand (Nice to Have)
- Performant distributed workload basics (Nice to Have)
- Job schedulers (Nice to Have)
- NCCL for performance troubleshooting (Nice to Have)
- Infrastructure as Code (Nice to Have)
- Config management tools (Ansible, Terraform) (Nice to Have)
- GitOps (Nice to Have)
- CI/CD participation (Nice to Have)
- GitHub Actions (Nice to Have)
- Access and security tooling (Teleport, Vault) (Nice to Have)
- Relevant certifications (Linux, Kubernetes, cloud, security) (Nice to Have)
Location
- West Coast Candidates Preferred
Work Type
- Onsite
- Travel
Experience Level
- 2-4 years experience
Salary/Compensations
- $100,000—$140,000 USD
Benefits
- Highly competitive package (base + equity)
- Reviews every 12 months
- Dynamic progression plan
- Human-First Flexibility
- Medical
- Dental
- Vision
- Flexible paid time off
- Parental leave
- Retirement plan participation
About the Company
- Nscale is the GPU cloud engineered for AI, providing cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers.
- Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development.
- Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.
- The Support and Operations team plays a critical role in maintaining service availability, driving service reliability and rapid response to customer tickets.
- We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency.
- As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work.
- Joining the team means contributing to building the technology that powers the future.
- Nscale offers a collaborative, supportive, and innovative environment where contributions spark real impact.
- Join the fastest-growing tech startup, push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI.
Equal Opportunity
- We strongly encourage applications from people of color, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, careers, and people from lower socio-economic backgrounds.
- If there’s anything we can do to accommodate your specific situation, please let us know.
