AI Infrastructure Engineer
Posted 1 day ago
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: AI Infrastructure Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $100,000–$160,000 Annually
Experience Required: 10+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary:
Key Responsibilities
- Design and operate GPU and accelerator infrastructure for training and inference, spanning on-prem clusters, cloud-managed services, and hybrid configurations.
- Build scheduling, queueing, and resource-sharing systems that maximize accelerator utilization across many teams.
- Integrate frameworks such as PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform offering.
- Operate high-performance storage systems and data pipelines that keep accelerators fed with training data at near-line-rate.
- Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth collective communication.
- Build observability for AI workloads including utilization, throughput, training stability, and failure-mode analytics.
- Implement checkpointing, restart, and fault-tolerance patterns for long-running training jobs at scale.
- Drive cost optimization across compute, storage, and networking through scheduling, spot capacity, and right-sizing.
- Develop developer tooling and paved-road workflows that let researchers launch experiments safely and efficiently.
- Partner with research and applied ML teams to plan capacity for upcoming training runs.
- Implement security controls, isolation, and access management for multi-tenant AI infrastructure.
- Drive automation across cluster provisioning, lifecycle management, and configuration enforcement.
- Maintain runbooks, capacity dashboards, and operational documentation for the AI platform.
- Stay current with AI infrastructure research, accelerator hardware, and emerging open-source AI tooling.
- Bachelor’s or Master’s degree in Computer Science or a related field.
- Ten or more years of experience in infrastructure, platform, or HPC engineering.
- Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
- Strong proficiency in Python and at least one systems language such as Go or C++.
- Deep understanding of distributed training, accelerator architectures, and collective communication.
- Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads.
- Strong understanding of Linux internals, networking, and high-performance storage.
- Experience with at least one major cloud provider’s ML infrastructure offerings.
- Strong software engineering practices including testing, CI/CD, and code review.
- Excellent communication and cross-functional collaboration skills.
- Experience operating InfiniBand or RDMA networking at scale.
- Contributions to open-source ML infrastructure projects.
- Familiarity with custom orchestrators or research-grade training stacks.
- Exposure to frontier model training operations.
- Experience with FinOps for AI workloads.
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3899. Learn more about Bright Vision Technologies at www.bvteck.com.
Bright Vision Technologies is an Equal Opportunity Employer.
Similar Infrastructure Engineer jobs(8)
Bright Vision Technologies is a minority-owned, product-focused AI technology company headquartered in Bridgewater, New Jersey. Founded in July 2020, we are a pure-play product engineering firm dedicated to transforming the staffing and enterprise automation industries through our flagship platform — Lumina. Lumina is an advanced AI-powered talent intelligence and enterprise automation platform that integrates artificial intelligence, hybrid cloud, and blockchain technologies to revolutionize how organizations source talent and automate complex workflows. What Lumina delivers: - AI-driven candidate sourcing, semantic resume matching, and intelligent ranking - Generative AI and Large Language Model orchestration with LangChain and RAG-based architecture - Automated recruiter assistant and workflow automation - Blockchain-enabled credential verification and tamper-proof digital identity - Advanced optimization for talent supply chain, workforce planning, and resource allocation - Enterprise-grade security with role-based access, encryption, and comprehensive audit logging - Seamless integration with CRM, ERP, HRMS, ATS, and legacy enterprise systems Built on a scalable microservices architecture with cloud-native deployment across AWS, Azure, GCP, and Red Hat OpenShift, Lumina is trusted by enterprises across financial services, healthcare, manufacturing, and retail — delivering efficiency improvements of 40–60% and measurable ROI within the first quarter of deployment. At Bright Vision, we are building the future of AI-powered talent intelligence — with a strong commitment to innovation, enterprise security, and U.S.-based job creation in AI engineering, product development, and go-to-market functions. Learn more at www.bvteck.com
Key team members
Jobr aggregates jobs directly from company career portals — no middlemen. Our team applies on your behalf with AI-tailored resumes, reviewed by a human before submission.
