Search

Artificial Intelligence Engineer LLM

PublishedPublished: 6/14/2022
Technology

Job Description

LLM Engineer

\n

Top Secret Clearance (TS/SCI) is required

\n

We are seeking an experienced LLMOps Engineer to deploy, operate, and optimize self-hosted Large Language Models (LLMs) in production.

\n

Understanding of how modern LLMs work, including transformer architectures, inference optimization, quantization, and fine-tuning, and hands-on experience managing production AI infrastructure.

\n

    \n
  • Deploy, manage, and optimize self-hosted LLMs using vLLM on GPU infrastructure.
  • \n

  • Configure multi-GPU deployments, including tensor parallelism, PagedAttention, continuous batching, and KV cache optimization.
  • \n

  • Deploy and serve open-weight models such as Llama, Mistral, Qwen, and other Hugging Face models.
  • \n

  • Optimize inference performance through quantization (AWQ, GPTQ, FP8) and benchmark latency, throughput, and GPU utilization.
  • \n

  • Build and maintain scalable Kubernetes-based AI infrastructure.
  • \n

  • Monitor model health, performance, and production reliability.
  • \n

  • Manage model versioning, deployments, rollbacks, and production releases.
  • \n

  • Implement CI/CD pipelines, Infrastructure-as-Code, observability, and incident response for AI platforms.
  • \n

  • Strong understanding of LLM architecture, transformer models, tokenization, inference, and fine-tuning.
  • \n

  • Hands-on experience deploying and operating vLLM in production.
  • \n

\n

Kubernetes, Docker, GPU clusters, and cloud platforms (AWS, Azure, or GCP).

\n

    \n
  • Open-source LLMs (Llama, Mistral, Qwen, etc.).
  • \n

  • Experience optimizing inference using tensor parallelism, PagedAttention, continuous batching, KV cache management, and quantization.
  • \n

  • Strong Python and Linux administration skills.
  • \n

  • Prometheus and Grafana.
  • \n

  • Experience with Infrastructure-as-Code (Terraform, Helm, Ansible) and CI/CD tools.
  • \n

\n

Preferred Skills

\n

    \n
  • Hugging Face
  • \n

  • LoRA fine-tuning
  • \n

  • RAG architectures
  • \n

  • Vector databases
  • \n

  • TensorRT-LLM
  • \n

  • Triton Inference Server
  • \n

  • Ray Serve
  • \n

  • NVIDIA CUDA, A100, and H100 GPUs
  • \n

  • LangChain
  • \n

  • DevSecOps and Site Reliability Engineering (SRE)
  • \n

\n

Keywords: SGLang LLMOps, MLOps, vLLM, Kubernetes, Docker, Python, Hugging Face, Llama, Mistral, Qwen, TensorRT-LLM, Triton, Ray Serve, CUDA, NVIDIA GPU, Terraform, Helm, AWS, Azure, GCP, Prometheus, Grafana, CI/CD, Infrastructure as Code, RAG, Vector Database.

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...