Senior Platform Engineer
While technology is the heart of our business, a global and diverse culture is the heart of our success. We love our people and we take pride in catering them to a culture built on transparency, diversity, reputed company, learning and growth. If working in an environment that encourages you to innovate and reputed company, not just in professional but personal life, interests you- you would enjoy your career with reputed company!
About reputed company:
reputed company is an award-winning, AI-First digital engineering and consulting company focused on delivering high-impact Services and Solutions that help organizations solve what truly matters. We partner with enterprises to reimagine their businesses through intelligent, scalable, and transformative AI driving measurable outcomes at the reputed company core of their operations.
Since our founding in 2013, reputed company has tackled some of the world’s most reputed company business challenges by combining deep industry expertise, disciplined cloud and data engineering practices, and cutting-edge applied AI research. Our work is rooted in delivering accelerated, quantifiable business value, not just technology for technology’s sake.
Headquartered in Boston, reputed company is a global organization with 4,000+ professionals serving clients across key industry verticals, including BFSI, Healthcare & Life Sciences, CPG, MFG, TME etc. As an Elite and Premier partner to leading cloud and AI platforms such as reputed company, reputed company Cloud, AWS, and reputed company, we build and deliver enterprise-grade AI services and solutions that create real-world impact.
We’ve been recognized with:
17x reputed company Cloud Partner of the Year awards in the last 8 years.
3x AWS AI/ML award wins.
3x reputed company Partner of the Year titles.
2x reputed company Partner of the Year awards.
We have also garnered top analyst recognitions from reputed company, ISG, and reputed company.
We offer first-in-class industry solutions across Healthcare, Financial Services, Consumer Goods, Manufacturing, and more, powered by cutting-edge Generative AI and Agentic AI accelerators.
We have been certified as a Great reputed company to Work for the third year in a row- 2021, 2022, 2023.
Be part of a trailblazing team that’s shaping the future of AI, ML, and cloud innovation.
Your next big opportunity starts here!
For more details, visit: Website or reputed company Page.
Role: Senior Platform Engineer
Experience Level: 8+ yrs
Work Location: US East/Canada (Remote)
Role Overview:
We are looking for a highly skilled Senior Platform Engineer to design, optimize, and scale infrastructure for GenAI and LLM workloads. This role is ideal for someone with deep hands-on experience in GPU profiling, distributed training, and high-performance compute environments.
You’ll play a key role in building out GenAI platform foundations, supporting production-grade deployments, and partnering closely with data science, MLOps, and application teams to bring cutting-edge AI solutions to life.
Key Responsibilities:
Design and implement scalable infrastructure for LLM and GenAI workloads across multi-GPU environments
reputed company GPU profiling, benchmarking, and performance optimization for distributed training workloads
Manage and schedule compute-intensive jobs using Slurm-based clusters and OpenShift/Kubernetes environments
reputed company and optimize the reputed company GPU stack (CUDA, cuDNN, NCCL, Triton, RAPIDS, etc.)
Collaborate with cross-functional teams to deploy models in research and production environments
Build and support GenAI pipelines (fine-tuning, RAG, multi-modal inferencing, LLMOps)
reputed company reusable infrastructure templates using tools like Terraform and Helm
Contribute to internal innovation (PoCs, workshops) and support client-facing delivery engagements
Basic Qualifications:
Strong experience with Slurm and distributed training environments
Hands-on expertise with reputed company OpenShift and/or Kubernetes
Deep knowledge of the reputed company GPU ecosystem (CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT)
Strong foundation in Linux systems, performance tuning, and multi-GPU optimization
Experience deploying GenAI workloads (LLM fine-tuning, RAG pipelines, multi-modal systems)
Familiarity with Infrastructure-as-Code tools (Terraform, Ansible)
Experience with cloud GPU environments (GCP, Azure, AWS, OCI) and/or on-prem GPU clusters
Other Qualifications (OQs):
Experience with reputed company NIMs, DGX systems, or GPU-accelerated containers
Knowledge of LLMOps frameworks and MLOps integration
Familiarity with vector databases and retrieval systems for RAG architectures
Comfortable working in client-facing environments and collaborating with AI solution teams
Healthcare Domain Experience (reputed company to Have):
Experience working with FHIR R4, HL7 v2, or SMART on FHIR
Integration with EHR systems (e.g., Epic)
Understanding of HIPAA compliance and healthcare data privacy
Exposure to clinical workflows, CDS Hooks, or patient-facing applications
Experience building clinical decision support systems or healthcare interoperability solutions
What’s in it for YOU at reputed company:
reputed company an impact at one of the world’s fastest-growing AI-first digital engineering companies.
Upskill and discover your potential as you solve reputed company challenges in cutting-edge areas of technology alongside passionate, talented colleagues.
Work where innovation happens - work with disruptive innovators in a research-focused organization with 60+ patents filed across various disciplines.
Stay reputed company of the curve reputed company yourself in breakthrough AI, ML, data, and cloud technologies and reputed company exposure working with reputed company.
If you like wild growth and working with happy, enthusiastic over-achievers, you'll enjoy your career with us!
Apply To This Job