About the Role
The Kotlin AI Value Stream team is responsible for how AI agents understand, generate, and improve Kotlin code across all platforms. This role will own the end-to-end loop of analyzing agent failures, building evals, researching fixes, and measuring improvements, directly shaping how millions of developers experience Kotlin through AI coding agents.
Responsibilities
- Build tools for agentic error analysis, including systematic capture, classification, and analysis of AI coding agent errors.
- Build observability pipelines over agentic traces, mining patterns from agent sessions in various coding agents.
- Build evaluation pipelines to measure Kotlin code generation quality across dimensions like correctness, idiomaticity, and test coverage.
- Build simulation environments for measuring coding agents on realistic Kotlin developer tasks.
- Own evaluation infrastructure, including metrics, experiment tracking, automated regression checks, and reproducible benchmarking.
- Research methods for improving agent and model behavior on Kotlin, experimenting with post-training techniques and context engineering approaches.
- Run experiments to measure impact using A/B comparisons, benchmark suites, and before/after analyses.
- Collaborate with model providers (Anthropic, OpenAI, and Google) to translate Kotlin-specific findings into model improvements.
- Design and build open-source benchmarks to measure AI coding agent performance on Kotlin tasks.
- Create task datasets covering the breadth of Kotlin usage.
- Maintain and evolve benchmarks as models improve, ensuring they remain challenging, relevant, and contamination-resistant.
Requirements
- Hands-on experience building evaluation or analysis pipelines for LLMs or AI coding agents in a research or production setting.
- Strong Python engineering skills (at least three years), with the ability to write clean, maintainable code.
- Experience with data analysis at scale: querying large datasets (SQL/Athena), building data pipelines, and performing statistical analysis of experimental results.
- Ability to own projects end to end – from problem identification to shipping a fix.
- A product-aware mindset, caring about how agents are used by developers and translating failure modes into evaluation and training work.
- Familiarity with Kotlin or a strong willingness to develop deep Kotlin expertise.
Skills
- Python
- Data analysis
- SQL
- Athena
- LLM evaluation
- AI coding agent analysis
- Post-training LLMs (SFT, DPO, GRPO)
- Deep learning frameworks (PyTorch)
- LLM training stacks (TRL, verl, Megatron)
- AI agent development
- Evaluation frameworks and tools (Inspect AI, Promptfoo, LM-evaluation-harness)
- Experiment tracking and observability (Weights & Biases, MLflow, Langfuse)
- Kotlin ecosystem (Android, Gradle, KMP, Spring, Ktor)
- Open-source contributions
Location
- Remote
Work Type
- Remote
- Work from home
- Work from office
Experience Level
- 3+ years of Python engineering experience
Salary/Compensations
- Strong base salary. We offer competitive pay that reflects your skills and experience.
Benefits
- Flexible work location
- Remote work (up to 30 days per year from abroad)
- Extra time off
- Medical insurance allowance
- Learning and development opportunities (conferences, courses, language classes)
- Relocation support
- Language classes
- Hot meal or lunch allowance on workdays
- Mental health support
- Sports benefit (on-site gym or sports club stipend)
- Internal events
About the Company
- At JetBrains, code is our passion. Ever since we started, back in 2000, we've been striving to make the strongest, most effective developer tools on earth.
Equal Opportunity
- We are an equal opportunity employer.
- We know great ideas can come from anyone, anywhere. That’s why we do our best to create an open and inclusive workplace – one that welcomes everyone regardless of their background, identity, religion, age, accessibility needs, or orientation.
- We process the data provided in your job application in accordance with the Recruitment Privacy Policy.
