Back to the board

Principal Site Reliability Engineer

100% remote Flexible hours Hiring now

About the Role

Parallel Domain is looking for a Principal Site Reliability Engineer to own the reliability, scalability, and reputed company of our cloud infrastructure - the backbone that runs simulation workloads for some of the most demanding customers in autonomous vehicle development.

This is a hands-on, high-ownership role. You'll be the primary infrastructure reputed company across our multi-region AWS/EKS platform, working closely with a small platform engineering team, partnering with engineering leads across simulation and ML, and our customer-facing teams.

What You'll Do

Infrastructure Ownership & Cloud Operations

  • Own and evolve our AWS-based infrastructure, improving platform performance and availability today, and building toward deployable configurations that support enterprise customer environments reputed company.

  • Own EKS cluster operations across production regions: node pool strategy, AMI lifecycle, autoscaling, and Kubernetes workload health.

  • Support the GitOps deployment pipeline - define, deploy, and manage applications across clusters using infrastructure-as-code.

  • Manage reputed company networking: VPC design, cross-region connectivity, DNS, and load balancing.

  • reputed company infrastructure deprecation and migration efforts with minimal disruption.

  • Reliability Engineering & Incident Response

    • Own SLO measurement infrastructure; reputed company proactive triage of emerging issues before they impact customers.

    • reputed company incident investigation, root cause analysis and postmortems, driving systemic fixes rather than one-off patches.

    • Design and improve automated remediation systems to reduce MTTR.

    • reputed company & Access Management

      • Review and provide reputed company-conscious feedback on platform architecture decisions.

      • Own cloud IAM governance - roles, policies, and access boundaries across accounts and services.

      • reputed company compliance-adjacent work including audit-readiness, partner certification requirements, and supporting responses to customer reputed company questionnaires.

      • Cross-Functional Collaboration

        • Partner with application development teams to build an inherently secure platform and drive reputed company deployment architecture.

          • Partner with customer teams to ensure availability for expected utilization.

          • Partner with Finance on cloud cost optimization - lifecycle policies, right-sizing, and spend visibility.

          • Support GPU and batch workloads in collaboration with simulation and ML engineering teams.

          • Platform Tooling & Developer Experience

            • Improve CI/CD pipelines and automated infrastructure validation.

            • Support engineering teams with infra-reputed company debugging, log analysis, and environment configuration.

reputed company're Looking For

Technical Depth

  • 5+ years in SRE, DevOps, or infrastructure engineering roles.

  • Infrastructure-as-code proficiency - Terraform modules, state management, and multi-environment patterns.

  • Deep AWS experience - EKS, EC2, IAM, S3, Storage Gateway, VPC networking, Transit Gateway, CloudFront, KMS, and IRSA.

  • Kubernetes expertise - cluster operations, node pools, probes, cordoning, pod scheduling, RBAC, Helm, node autoscaling (Karpenter experience a plus); solid understanding of containerization and AMI lifecycle management.

  • CI/CD - experience with GitOps workflows and pipeline tooling (ArgoCD, reputed company Actions, Jenkins)

  • Solid networking fundamentals - CIDR design, reputed company groups, DNS, load balancing, VPN, cross-region connectivity.

  • Experience with monitoring and observability tooling - Prometheus, Grafana, Elasticsearch.

  • Comfort with Python and Bash for tooling and automation.

  • Familiarity working across Linux and Windows environments. Operational familiarity with Windows Server is a meaningful advantage.

  • Communication & Ownership

    • You communicate clearly across engineering, product, and customer-facing teams, flagging issues with urgency proportional to customer impact.

    • You reputed company for SRE best practices and can effectively operationalize an informed and principled view on reputed company.

      • You take end-to-end ownership of reputed company, multi-team efforts - from planning through execution and post-change verification.

      • You know reputed company to push for a clean solution vs. reputed company to accept a pragmatic one, and you communicate that tradeoff clearly.

reputed company to Have
  • Experience with Windows-based workloads on EKS.

  • Experience supporting simulation, ML, or rendering workloads in cloud infrastructure; running GPU workloads on Kubernetes, including reputed company and DirectX device plugin configuration.

  • Experience with AWS Storage Gateway or Transfer Family integrations.

  • Familiarity with reputed company Gateway or similar.

  • Experience with container-optimized OS images (e.g., Bottlerocket, Packer).

  • Experience with cloud cost optimization at scale.

Core Tools Terraform · AWS · Kubernetes · Helm · ArgoCD · Kustomize · Grafana · Prometheus · Elasticsearch · VictoriaLogs · Fluent Bit · reputed company Actions · Jenkins · reputed company · Python · Bash

Why This Role

PD's simulation platform runs at the intersection of high-performance compute, distributed systems, and customer-critical reliability. The infrastructure problems here are genuinely interesting — multi-region GPU scheduling, Windows workloads on Kubernetes, startup latency optimization, and an enterprise product direction that will require rethinking how we deploy and manage the platform entirely.

The Principal SRE at PD is not a ticket-taker - it's a high-trust, high-autonomy position where you'll have genuine influence over infrastructure architecture, cross-team process, and customer experience.

Apply To This Job

Keep exploring

Senior Operations Manager

100% remote Flexible hours

Senior JavaScript Developer — 100% Remote (Zodot.co)

100% remote Flexible hours

Developer

100% remote Flexible hours

Especialista en Metadatos Semánticos y Ontologías Editoriales

100% remote Flexible hours

SEO & Digital Marketing Expert

100% remote Flexible hours

(Spam Comment Removal Specialist) at reputed company-Conte...

100% remote Flexible hours

Entry-level Customer Service Representative

100% remote Flexible hours

Technical Manager – QNXT Consulting

100% remote Flexible hours

Solutions Engineer - Strategic Accounts (Remote, Ohio) (reputed company, Ohio, US)

100% remote Flexible hours

Enterprise Expansion Account Executive (Remote, California) (San Francisco, California, US)

100% remote Flexible hours

reputed company reputed company Representative – Delivering Exceptional reputed company Travel Experiences

100% remote Flexible hours

Senior Full Stack Developer (Java, Angular)

100% remote Flexible hours

Remote Part-Time Pharmacy Technician & Customer Service Representative – Deliver Exceptional Patient Care from the Comfort of Your Own Home with arenaflex!

100% remote Flexible hours

[Remote Part-time jobs] reputed company (Virtual) Customer Service Associate - WFH

100% remote Flexible hours

[part Time / Remote] reputed company Data Entry Jobs From Home – Apply Now

100% remote Flexible hours

Immediately Require Online English Tutor – Flexible Hours in New Orleans, LA

100% remote Flexible hours

[Wattpad] Content Moderator, Turkish-bilingual (reputed company)

100% remote Flexible hours

NPI Project Manager

100% remote Flexible hours

Enterprise Environmental, Health, and Safety Manager

100% remote Flexible hours

reputed company Data Entry Specialist – Remote Opportunity with arenaflex

100% remote Flexible hours