Job Description
Job Title/Role: Senior Generative AI Architect
\n
Key Skills: Python, GenAI & LLM Engineering,
\n
Experience: 10-15 Years experience
\n
Location: Atlanta, GA
\n
\n
We at Coforge are hiring an experienced Senior Generative AI Architect with strong expertise in Generative AI, cloud-native architecture. The ideal candidate will have a hands-on background in Generative AI solutions, system design, and AI-driven customer experience solutions, with the ability to design, build, and scale intelligent solutions for enterprise clients.
\n
\n
Key Responsibilities:
\n
Model Fine-Tuning & Core ML Expertise:
\n
- \n
- Fine-Tuning Experience.
- End-to-end process followed: data preparation → training → validation → deployment.
- Dataset curation strategies (cleaning, labeling, augmentation, handling noise).
- Iterative training approach: number of cycles, convergence criteria, and evaluation metrics.
\n
\n
\n
\n
\n
Model Design & Architecture Decisions.
\n
- \n
- Rationale for selecting model architecture (Transformer vs. classical ML approaches).
- Understanding of pre-trained models vs. custom models.
- Layer-level customization (e.g., freezing/unfreezing layers, adapter layers, LoRA).
\n
\n
\n
\n
Optimization & Training Techniques.
\n
- \n
- Choice of loss functions and their business/technical rationale.
- Gradient-related challenges (vanishing/exploding gradients) and mitigation techniques.
- Handling imbalanced datasets (resampling, weighting, synthetic data generation).
\n
\n
\n
\n
Model Performance & Stability.
\n
- \n
- Ranking mechanisms (e.g., low-rank adaptations, embedding ranking logic).
- Managing model drift (data drift, concept drift detection and remediation strategies).
\n
\n
\n
Post-Training Strategy.
\n
- \n
- Model evaluation, monitoring, and retraining pipelines.
- Observability and feedback loops (model metrics, user feedback integration).
- Deployment validation and A/B testing approaches.
\n
\n
\n
\n
\n
Agentic AI & LLM Application Design:
\n
Prompting Techniques.
\n
- \n
- Zero-shot vs. few-shot prompting strategies and when to use each.
- Prompt engineering and prompt fine-tuning techniques.
\n
\n
\n
Frameworks & Libraries.
\n
- \n
- Experience with agentic AI frameworks (e.g., LangChain, Semantic Kernel, AutoGen, CrewAI).
- Integration patterns for tool usage and orchestration.
\n
\n
\n
Context Engineering.
\n
- \n
- Techniques to manage context windows effectively.
- Retrieval-Augmented Generation (RAG) design and optimization.
\n
\n
\n
Token Economy Optimization.
\n
- \n
- Cost optimization strategies (prompt compression, chunking, caching).
- Trade-offs between latency, cost, and accuracy.
\n
\n
\n
Agent Architecture.
\n
- \n
- Design of self-healing systems (retry logic, fallback strategies, tool re-planning).
- Memory management (short-term vs. long-term; local vs. global memory).
- Best practices in agent orchestration and modular design.
\n
\n
\n
\n
Codebase & Project Structure.
\n
- \n
- Ideal structure for scalable AI/agentic applications.
- Separation of concerns (prompts, tools, memory, orchestration layers).
\n
\n
\n
\n
LLM Observability & Data Architecture:
\n
LLM Observability.
\n
- \n
- Monitoring LLM outputs (quality, hallucination detection, latency, cost).
- Instrumentation and logging strategies.
\n
\n
\n
Vector Databases & Retrieval.
\n
- \n
- Experience with vector DBs (e.g., Pinecone, FAISS, Weaviate, Azure AI Search).
- Embedding strategies and indexing mechanisms.
- Distance/similarity metrics (cosine similarity, Euclidean, dot product) and use cases.
\n
\n
\n
\n
Hybrid Data Architectures.
\n
Combining Vector DBs with:
\n
- \n
- Graph DBs (relationship-driven queries).
- Relational DBs (structured data).
- Metadata stores (filtering, search refinement).
- Designing efficient retrieval pipelines.
\n
\n
\n
\n
\n
\n
MCP (Model/Modular Control Plane / Model Context Protocol) & Deployment:
\n
MCP / Orchestration Layer.
\n
- \n
- Understanding and application of MCP concepts in AI systems.
- Managing communication between models, tools, and services.
\n
\n
\n
Deployment Strategies.
\n
- \n
- Model deployment patterns (batch, real-time, streaming).
- Containerization (Docker/Kubernetes) and cloud deployment (Azure/AWS/GCP).
- CI/CD pipelines for AI models.
\n
\n
\n
\n
Scalability & Reliability.
\n
- \n
- Load handling, auto-scaling, and failover mechanisms.
- Performance optimization in production environments.
\n
\n
\n
\n
Good to Have:
\n
- \n
- Designing test strategies for deterministic and non-deterministic (AI) systems.
- Establishing measurable benchmarks for LLM performance and system reliability.
\n
\n
\n
