Inscreva-se para aceder a todos os recursos do nosso serviço
  • Pesquisa de emprego
  • Favorito
  • Criar um CV
    Novo
  • Salários
  • Alertas de emprego

Technical Lead - GPU Infrastructure

Portugal
  • Trabalho remoto

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Technical Lead - GPU Infrastructure based in Portugal.

This is a hands-on technical leadership role responsible for architecting and delivering a large-scale GPU infrastructure platform.
You will lead the evolution from managed Kubernetes workloads toward bare-metal GPU infrastructure, including Slurm-based research computing and Kubernetes-powered inference.
The role combines deep systems expertise with engineering leadership, team management, and direct ownership of architecture and delivery.
You will oversee a distributed team spanning backend, frontend, DevOps, QA, and documentation while maintaining high technical standards.
Your work will support research, model training, and managed inference workloads requiring reliable, scalable, and observable GPU compute.
You will also serve as the primary technical interface with infrastructure partners, hardware providers, and internal platform consumers.
This is an opportunity to shape the architecture and operational foundations of a sophisticated GPU platform in a fully remote environment.

Accountabilities:

The Technical Lead will own the platform architecture and engineering delivery while remaining deeply involved in technical decisions, infrastructure operations, team leadership, and partner relationships.

  • Own the end-to-end platform architecture, including architecture proposals, high-level and low-level designs, technical reviews, and ongoing architecture documentation.
  • Lead and line-manage a distributed engineering team across backend, frontend, DevOps, QA, and documentation.
  • Establish engineering standards, oversee code and design reviews, manage release gates, conduct one-to-ones, and provide growth and performance feedback.
  • Design, build, and operate a managed Slurm service supporting research and model-training workloads.
  • Own Slurm controllers, accounting, partitions, login nodes, node onboarding, acceptance testing, driver and CUDA baselines, upgrades, stalled-job detection, node health, draining, autohealing, storage visibility, identity, and workload isolation.
  • Lead GPU infrastructure operations on bare-metal environments, including NVIDIA drivers, CUDA, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance processes.
  • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal, including NVIDIA GPU Operator and Network Operator.
  • Oversee GPU isolation using technologies such as KubeVirt and VFIO and manage day-two infrastructure operations, upgrades, backup, recovery, and node replacement.
  • Define managed inference architecture covering serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute capabilities.
  • Establish observability across the control plane, GPU fleet, and application layers through metrics, logging, alerting, and SLOs.
  • Lead incident response, post-incident reviews, and the development of an on-call model that is sustainable for a lean engineering organization.
  • Act as the primary technical interface with infrastructure partners and vendors, translating requirements into written specifications and acceptance tests.
  • Manage technical escalations with partners through resolution and contribute to capacity planning and hardware sourcing decisions.
  • Work directly with research, model-training, and product teams to translate workloads into platform requirements and manage capacity constraints.
  • Hire and develop members of the platform team while maintaining a high technical bar.
  • Contribute to architecture decisions involving distributed systems, high-performance computing, networking, storage, virtualization, and GPU workloads.
  • Requirements:

    The ideal candidate combines deep hands-on GPU and infrastructure expertise with proven technical leadership experience. They should be comfortable operating complex production systems, making architecture decisions, leading distributed teams, and remaining close to the code and infrastructure.

    • 8+ years of hands-on engineering experience, including at least 3 years leading teams responsible for infrastructure platforms used by other teams.
    • Bachelor's or Master's degree in computer science, engineering, or a related field, or equivalent practical experience.
    • Extensive hands-on experience operating Slurm in production, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health, and upgrades.
    • Experience operating HPC or GPU training clusters for research or model-development users.
    • Strong experience operating NVIDIA GPU fleets on bare metal, including driver and CUDA lifecycles, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance.
    • Deep knowledge of InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance issues.
    • Strong Linux systems expertise, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning.
    • Proven production Kubernetes experience covering control planes, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy.
    • Experience with HPC storage and large-scale data movement, including shared filesystems such as VAST, Lustre, or NFS and node-local NVMe caching.
    • Experience distributing large model weights and datasets across multiple nodes.
    • Strong observability and operations experience with Prometheus, Grafana, Loki, or comparable platforms, including SLOs, incident response, and post-incident reviews.
    • Working proficiency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and make architecture decisions.
    • Experience delivering a multi-tenant IaaS, PaaS, research computing service, or comparable platform with resource isolation, quotas, usage metering, APIs, and CLI interfaces.
    • Demonstrated people leadership across time zones and the ability to lead cross-functional technical reviews.
    • Strong written architecture and decision-making skills, including documenting alternatives and trade-offs.
    • Confidence communicating technical decisions and respectfully challenging partners or executives when necessary.
    • Excellent written and spoken English.
    • Fully remote availability with a working location between UTC and UTC+5:30 to provide overlap with teams and partners across Europe and India.
    • Willingness to travel occasionally to partner sites and team events.
    • Desirable experience includes:

      • Slurm operators on Kubernetes, such as Soperator or Slinky, or Kubernetes-native schedulers such as Kueue, Volcano, KAI, or Kubeflow Trainer.
      • Modern model-serving technologies such as vLLM, SGLang, or TensorRT-LLM.
      • GPU parallelism strategies, quantization trade-offs, and GPU memory planning.
      • Multi-tenant GPU isolation using KubeVirt, Kata Containers, QEMU/KVM, Firecracker, or similar technologies.
      • Confidential computing technologies such as Intel TDX, AMD SEV-SNP, or NVIDIA confidential-compute capabilities.
      • Cluster API, kubeadm, Cilium, GPU autohealing, infrastructure as code, and GitOps.
      • Experience working on the operator side of a GPU cloud, university or national HPC center, or AI research platform.
      • Peer-to-peer or distributed-systems experience.
      • Experience working with hardware providers responsible for provisioning but not operating infrastructure, including establishing contracts and acceptance tests.
      • Benefits:

        • 100% remote position.
        • Opportunity to lead the architecture and delivery of a sophisticated GPU infrastructure platform.
        • High-impact technical leadership role spanning bare-metal GPU infrastructure, Slurm, Kubernetes, inference, and observability.
        • Leadership responsibility for a distributed engineering organization across multiple technical disciplines.
        • Significant ownership over architecture, engineering standards, delivery planning, and team development.
        • Direct involvement with infrastructure partners and hardware providers.
        • Opportunity to support advanced AI research, model training, and managed inference workloads.
        • International and distributed working environment with colleagues and partners across Europe and India.
        • Occasional opportunities for travel to partner locations and team events.
        • Opportunity to work at the intersection of high-performance computing, AI infrastructure, distributed systems, and cloud-native technologies.
Vaga publicada 18 horas atrás
Empregos semelhantes que podem ser interessantes para vocêCom base na vaga Technical Lead - GPU Infrastructure em Portugal
  •  ...partner is looking for a Senior Technical Program Manager based in...  ...partner closely with engineering leads and product managers to...  ...external commitments are made. ~ Lead complex programs end-to-end,...  ...within AI/ML, platform infrastructure, developer tooling, or similarly... 
    Trabalho remoto
    Portugal
    dia atrás
  •  ...next steps. Our partner is looking for a Technical Delivery Manager based in Portugal. This...  ...for an experienced technology delivery leader who can balance people, technology, execution...  ...dependencies, and on-time outcomes. ~ Lead delivery rhythms covering sprint planning... 
    Trabalho remoto
    Portugal
    18 horas atrás
  •  ...all applications and next steps. Our partner is looking for a Technical Business Analyst based in Portugal. This is a remote contract...  ...communication and organizational discipline. Accountabilities ~ Lead structured discovery workshops with engineering, product,... 
    Trabalho remoto
    Portugal
    dia atrás
  •  ...We're looking for an SAP MDG Technical Architect to join our team in Portugal in a remote working mode. In this role, you will lead the technical design, implementation, and integration of SAP MDG solutions for material master governance. You will bridge functional requirements... 
    Trabalho remoto

    EPAM Systems

    Portugal
    26 dias atrás
  •  ...global semiconductor manufacturing and high-tech innovator. As a Lead S/4 HANA Developer , you will shape data models, ETL pipelines...  ...and retrospectives while delivering incremental value Document technical standards, best practices, and processes to support... 
    Trabalho remoto

    EPAM Systems

    Portugal
    8 dias atrás
  •  ...global Intelligent Support-as-a-Service leader, partnering with tech companies and industry...  ...2010 to deliver secure customer and technical support. We operate globally, supporting...  ...had a chance to be a part of the world's leading SaaS, software, or hardware solutions?... 

    SupportYourApp

    Portugal
    21 dias atrás
  •  ...Do you speak German? &##128640; A new opportunity in Technical Support for EV Charging (Electric Vehicles) is opening a new chapter in...  ...Our client has been powering growth for disruptive brands and leading companies in the US and Europe since 2014. Their team is made up... 

    Speakit Jobs

    Portugal
    2 meses atrás
  •  ...Freelance Technical Trainer - Agricultural & Construction Machinery ABOUT THE OPPORTUNITY A global market leader in construction and agricultural machinery is looking for freelance technical trainers for a high-impact project. If you have solid experience with... 

    rpc - The Retail Performance Company

    Portugal
    2 meses atrás
  •  ...Become the technical expert who helps clients solve every problem! Work on-site in Lisbon providing high-level technical support and collaborate with international teams to ensure the best customer experience. Responsibilities: Troubleshoot and resolve technical... 

    Speakit Jobs

    Portugal
    2 meses atrás
  •  ...As a Technical Support Specialist, you will be the first point of contact for clients experiencing technical issues. Your role combines problem-solving, customer service, and technical knowledge to ensure fast and effective resolutions. This is ideal for someone who enjoys... 

    Speakit Jobs

    Portugal
    2 meses atrás
  •  ...requests. Provision of 1st line support for incidents. Responsibilities: Works under supervision, supporting standard technical queries related to a single product/small set of products (e.g. Microsoft products, operating system, basic networking, PCs).... 

    Speakit Jobs

    Portugal
    2 meses atrás
  •  ...Are you an experienced sales leader with a track record in the AdNetwork industry? We are seeking a Publisher Team Leader  (remote) to...  ...analytical skills ~ Fluent English Responsibilities: Lead and motivate the Publisher Manager team to achieve KPIs and growth... 

    Centro.team

    Portugal
    12 dias atrás
  • €280 por mês

     ...If you’re looking to grow and be inspired, as a Technical Support Digital Ads in Lisbon, Portugal (On-Site) you will make use of your skills, supporting our client's different areas of social media. And be an active part in ensuring client guidelines, promoting, and actively... 

    Speakit Jobs

    Portugal
    2 meses atrás
  •  ...for what’s next? Our client is a global technology and services leader that powers the brands of the future. They help well-known...  ...countries. If you’re looking to grow and be inspired, as a Customer Technical Advisor & Inbound Sales in Portugal, you will ensure that... 

    Speakit Jobs

    Portugal
    2 meses atrás
  • $1,250 por mês

     ...listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an Estimating & Design Lead - US Commercial Landscaping based in Portugal. This is a remote opportunity for an experienced landscaping professional who can... 
    Trabalho remoto
    Portugal
    dia atrás
  • €25,000 a €38,000 por año

     ...and affiliations. It will be available in 100+ countries by the end of 2025 and 125 countries by the end of 2025. As our associate lead, you will guide and develop a remote team of researchers who identify various data related to OpenData while following compliance... 
    Trabalho remoto

    Veeva Systems

    Portugal
    11 horas atrás
  •  ...We're looking for a Senior Platform Engineer — Agent Gateway Lead to join our team in Portugal in a remote working mode. In this role, you will lead the design and implementation of the AWS AgentCore Gateway for an enterprise AI agent development platform. You will... 
    Trabalho remoto

    EPAM Systems

    Portugal
    19 dias atrás
  • €1,250 a €1,400 por mês

     ...technology companies? We are looking for a Dutch speaking Technical Support to join a global leader in consumer electronics and home appliance. Join a...  ...Offered Accommodation  Excellent work opportunity in a leading multinational company Private health insurance... 

    Speakit Jobs

    Portugal
    2 meses atrás
  • €25,000 a €38,000 por año

     ...industry identifiers) and affiliations. It is available in 100+ countries today and growing. Your mission as our Data Operations Lead is to be the operational backbone of our curation efforts. You will guide and lead a remote team of data curators, ensuring we unlock... 
    Trabalho remoto

    Veeva Systems

    Portugal
    11 horas atrás
  • €25,000 a €38,000 por año

     ...employees, and communities. The Role OpenData Clinical is global reference data on sites and investigators. As our Associate Lead, you will lead a remote team of 30-50 freelancers who do rule-based web research and identify investigators and sites related to... 
    Trabalho remoto

    Veeva Systems

    Portugal
    11 horas atrás
  • We are looking for an experienced and performance-driven Team Lead to lead our Portuguese-language sales team. Our most profitable and strategically important market. This is a hands-on leadership role where you'll take full ownership of your team's commercial performance... 

    Best Service Team

    Portugal
    2 meses atrás
  • We are looking for an experienced SDR Team Lead who can build, coach, and scale a high-performing outbound sales team. This is not simply a people management role. We need someone who has successfully grown an SDR function, introduced new outbound strategies, and knows... 

    Bnberry

    Portugal
    2 meses atrás
  • €30,000 a €45,000 por año

     ...data applications to improve research and patient outcomes, powered by a global network of freelancers. As Data Operations Hiring Lead, you will develop a seamless hiring engine and lead a remote team of coordinators to turn recruitment into a smooth, end-to-end... 
    Trabalho remoto

    Veeva Systems

    Portugal
    11 horas atrás
  •  ...thinking, requiring you to care equally about technical quality and customer outcomes. You will...  ...and model pricing, inference economics, GPU costs, and the allocation of usage...  ...management products. ~ Knowledge of AI infrastructure, cloud economics, or AI workload cost optimization... 
    Trabalho remoto
    Portugal
    dia atrás
  •  ...You will define architecture strategy, technical roadmaps, and reference patterns while...  ...role combines deep expertise in cloud infrastructure, distributed systems, DevOps, automation...  ...approaches. ~ Mentor engineers and technical leads, identify capability gaps, promote... 
    Trabalho remoto
    Portugal
    18 horas atrás
  •  ...relationship management, commercial development, technical coordination, and operational leadership...  ...’ll work closely with engineering and infrastructure teams to resolve issues and ensure...  ...business development opportunities. Lead client communication during technical incidents... 
    Trabalho remoto
    Portugal
    18 horas atrás
  •  ...seamlessly with research, engineering, and product teams to ensure technical milestones translate directly into tangible user value....  ...Machine Learning role Experience with Python, PyTorch / JAX, GPU-based training and inference systems Demonstrated experience... 
    Trabalho remoto

    Sowelo Consulting sp. z o.o.

    Portugal
    7 dias atrás
  •  ...which has grown into an industry-leading, multibillion-dollar...  ...looking for a Dev Manager to lead the engineering team responsible...  ...engineering leadership with hands-on technical expertise to help customers...  ...access controls, and shared infrastructure. Own automated pipelines... 
    Trabalho remoto
    Portugal
    9 horas atrás
  •  ...engineering, and we’re looking for a seasoned architect to lead our client engagements and technical teams in Portugal. What We’re Working On -...  ...foundation in software engineering, MLOps, and cloud infrastructure (AWS/GCP). You know how to design for scalability, reliability... 

    Tensorops

    Portugal
    14 dias atrás
  •  ...and production AI engineering. You will work with leading foundation models and modern AI infrastructure to create reliable, observable, and low-latency...  ...first team, you will have significant ownership over technical architecture and product direction. This is an opportunity... 
    Trabalho remoto
    Portugal
    18 horas atrás

Deseja receber mais vagas?

Assine e receba vagas semelhantes a Technical Lead - GPU Infrastructure. Seja o primeiro a se candidatar!