About the Role
We are seeking a consultant with expertise in on-premise Large Language Model (LLM) and Vector Database implementation. This role involves deploying open-source LLMs, building RAG pipelines, and ensuring security and governance for enterprise environments.
Responsibilities
- Deploy open-source LLMs such as Meta Llama 3 and Mistral / Mixtral in on-prem or private environments
- Implement Retrieval-Augmented Generation (RAG) pipelines
- Generate and manage embeddings and metadata filtering
- Implement access controls and audit logging
- Provide reference architecture and deployment guidance
- Develop a working prototype (LLM + vector DB + RAG)
- Conduct documentation and knowledge transfer to internal teams
Requirements
- Hands-on experience deploying open-source LLMs in on-prem or private environments
- Strong proficiency in Python for LLM inference, prompt engineering, and integration
- Experience with CPU-based inference, model quantization, and performance tuning
- Practical experience with open-source vector databases
- Proven implementation of Retrieval-Augmented Generation (RAG) pipelines
- Experience generating and managing embeddings and metadata filtering
- Understanding of data privacy, air-gapped deployments, and enterprise security requirements
- Experience implementing access controls and audit logging
Skills
- Meta Llama 3
- Mistral / Mixtral
- Python
- Prompt Engineering
- CPU-based inference
- Model quantization
- Performance tuning
- Qdrant
- Chroma
- Milvus
- pgvector
- Retrieval-Augmented Generation (RAG)
- Embeddings
- Metadata filtering
- Data privacy
- Air-gapped deployments
- Enterprise security
- Access controls
- Audit logging
- LangChain
- LlamaIndex
- Rust
- Go
- C++
- Docker
- Kubernetes
- vLLM
- llama.cpp
- Hugging Face Transformers
Location
- Philadelphia
Work Type
- Contract
- Hybrid
Experience Level
- Consultant
