Jobiglo

Ni rezultatov.

Senior AI Compute Infrastructure Engineer

Kraken

Senior 🇬🇧 English
accelerator infrastructure node configuration scheduling primitives workload isolation quota management orchestration vLLM Triton Inference Server TensorRT observability for GPU utilization

Opis delovnega mesta

About the role

As a Senior AI Compute Infrastructure Engineer you will join Kraken’s dedicated AI Compute and Infrastructure team. You will be responsible for designing, operating, and scaling the GPU and accelerator clusters that power model training, inference, evaluation, and experimentation across the exchange.

Key responsibilities

  • Own and operate GPU/accelerator clusters, including drivers, runtimes, kernels, device plugins, and node configuration.
  • Design scheduling, orchestration, placement, quota management, and utilization systems for heterogeneous accelerator environments.
  • Optimize inference pipelines for latency, throughput, reliability, memory efficiency, and cost using frameworks such as vLLM, Triton Inference Server, or TensorRT.
  • Partner with ML engineers and researchers to remove bottlenecks in training, batch and online inference, deployment, and production debugging.
  • Build observability for GPU utilization, memory pressure, queue depth, token throughput, request latency, failed workloads, capacity pressure, and spend.
  • Drive reliability through incident response, alerting, runbooks, and post‑incident improvements for always‑on AI compute.

Required profile

  • Several years of experience operating large‑scale GPU or accelerator infrastructure.
  • Proven ability to work with AI/ML researchers, platform engineers, security, and product teams.
  • Strong problem‑solving mindset with focus on performance, cost efficiency, and reliability.

Required skills

  • GPU and accelerator cluster management
  • Drivers, runtimes, kernels, device plugins
  • Scheduling primitives, workload isolation, quota management
  • Orchestration and placement systems for heterogeneous hardware
  • Inference optimization using vLLM, Triton Inference Server, TensorRT
  • Observability tooling for GPU utilization, memory pressure, queue depth, token throughput, latency, and cost metrics

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Kraken.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Zakaj prijavljate to ponudbo?

Hvala za vaše sporočilo. Pregledali bomo to ponudbo.

Explore further

Salaries, guides and searches for Slovenija.

Prijavite se v 30 sekundah

Vnesite svoj e‑mail za prijavo. Račun bo ustvarjen samodejno.

Prijavi se zdaj →

Z nadaljevanjem sprejmete naše pogoje uporabe.

Že imate račun? Prijava

Une question sur cette offre ?

Posez-la ici : vous recevrez le récapitulatif de l'offre par e-mail, tout de suite.

💬 Chat with us on Telegram

Objavljeno pred 1 mesecem

Poteče čez 2 tedna

35 ogledi · 0 interested

Izboljšajte svoje možnosti

Naložite svoj življenjepis: predlagali bomo ponudbe, ki ustrezajo vašemu profilu.

Analiza vašega življenjepisa poteka...

Kraken