Artificial Intelligence Engineer LLM
Job Description
LLM Engineer
\n
Top Secret Clearance (TS/SCI) is required
\n
We are seeking an experienced LLMOps Engineer to deploy, operate, and optimize self-hosted Large Language Models (LLMs) in production.
\n
Understanding of how modern LLMs work, including transformer architectures, inference optimization, quantization, and fine-tuning, and hands-on experience managing production AI infrastructure.
\n
- \n
- Deploy, manage, and optimize self-hosted LLMs using vLLM on GPU infrastructure.
- Configure multi-GPU deployments, including tensor parallelism, PagedAttention, continuous batching, and KV cache optimization.
- Deploy and serve open-weight models such as Llama, Mistral, Qwen, and other Hugging Face models.
- Optimize inference performance through quantization (AWQ, GPTQ, FP8) and benchmark latency, throughput, and GPU utilization.
- Build and maintain scalable Kubernetes-based AI infrastructure.
- Monitor model health, performance, and production reliability.
- Manage model versioning, deployments, rollbacks, and production releases.
- Implement CI/CD pipelines, Infrastructure-as-Code, observability, and incident response for AI platforms.
- Strong understanding of LLM architecture, transformer models, tokenization, inference, and fine-tuning.
- Hands-on experience deploying and operating vLLM in production.
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
Kubernetes, Docker, GPU clusters, and cloud platforms (AWS, Azure, or GCP).
\n
- \n
- Open-source LLMs (Llama, Mistral, Qwen, etc.).
- Experience optimizing inference using tensor parallelism, PagedAttention, continuous batching, KV cache management, and quantization.
- Strong Python and Linux administration skills.
- Prometheus and Grafana.
- Experience with Infrastructure-as-Code (Terraform, Helm, Ansible) and CI/CD tools.
\n
\n
\n
\n
\n
\n
Preferred Skills
\n
- \n
- Hugging Face
- LoRA fine-tuning
- RAG architectures
- Vector databases
- TensorRT-LLM
- Triton Inference Server
- Ray Serve
- NVIDIA CUDA, A100, and H100 GPUs
- LangChain
- DevSecOps and Site Reliability Engineering (SRE)
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
Keywords: SGLang LLMOps, MLOps, vLLM, Kubernetes, Docker, Python, Hugging Face, Llama, Mistral, Qwen, TensorRT-LLM, Triton, Ray Serve, CUDA, NVIDIA GPU, Terraform, Helm, AWS, Azure, GCP, Prometheus, Grafana, CI/CD, Infrastructure as Code, RAG, Vector Database.
