About the Role
DV Trading is establishing a centralized AI function and is seeking to hire for the model layer. The goal is for DV to own its model capability, rather than relying on external providers. This role involves fine-tuning and distilling open-weight models for DV-specific tasks, operating on-prem inference infrastructure, and building a model gateway for intelligent routing across providers. The aim is to achieve lower cost and better latency in the near term, and a firm that controls its own AI stack in the long term.
Responsibilities
- Build and operate a model gateway routing inference across open and closed models with cost, latency, and quality tracking
- Design and run distillation pipelines: use frontier model outputs to generate training data for task-specific open models
- Fine-tune and evaluate open-weight models (Llama, Qwen, Mistral, or similar) for DV-specific tasks
- Deploy and maintain on-prem inference infrastructure (vLLM, TGI, or equivalent) on Kubernetes
- Build model evaluation frameworks for quality, cost, latency, and regression
- Define criteria and tooling for model selection: when open models are production-ready vs. when to use closed APIs
- Partner with the agent engineering team to ensure the model layer meets agent workload
Requirements
- 5+ years software engineering; strong Python
- Production fine-tuning or distillation of open-weight models (not just inference API wrappers)
- Experience serving LLMs on-prem (vLLM, TGI, Triton, or equivalent)
- Experience managing GPU infrastructure (provisioning, scheduling, utilization monitoring) in a production environment
- Model evaluation and regression testing in production
- Kubernetes and GPU workload management
- Strong grasp of the tradeoffs between open and closed models across cost, quality, latency, and data sensitivity
- Quantization, PEFT/LoRA, or other efficient training techniques
- Model gateway or inference proxy design (routing, fallback, rate limiting)
- Financial services or other regulated/sensitive-data environments
- Familiarity with the open model ecosystem (Hugging Face, model cards, licensing)
Skills
- Python
- Fine-tuning open-weight models
- Distillation pipelines
- LLM serving on-prem
- GPU infrastructure management
- Kubernetes
- Model evaluation
- Regression testing
- Quantization
- PEFT/LoRA
- Model gateway design
- Inference proxy design
- Hugging Face
- Model cards
- Licensing
Location
- Chicago
Work Type
- On-prem
Experience Level
- 5+ years software engineering
Salary/Compensations
- $200,000—$300,000 USD
Benefits
- Discretionary bonus eligibility
- Medical, dental, and vision insurance
- HSA, FSA, and Dependent Care Options
- Employer Paid Group Term Life and AD&D insurance
- Voluntary LTD, Life & AD&D insurance
- Flexible Vacation policy
- Retirement plan with employer match
About the Company
- Founded 20 years ago and headquartered in Chicago, the DV Group of financial services firms has grown to more than 600 people operating throughout North America, Europe and Asia.
- Since spinning out of a large brokerage firm in 2016, DV Trading has rapidly scaled as an independent proprietary trading firm utilizing its own capital, trading strategies, and risk management methodologies to provide liquidity to worldwide financial markets and hedging opportunities to commodity producers and users.
- DV group affiliates include two broker dealers, a cryptocurrency market making firm, and a bourgeoning investment adviser.
Equal Opportunity
- DV is proud to be an equal opportunity employer and committed to creating an inclusive environment for all employees.
