About the Role
Lead the design, build, test, and deployment of end-to-end AI solutions for complex document understanding tasks in the legal domain. Direct the execution of large-scale projects including advanced semantic chunking models, document enrichment systems, LLM-based knowledge graph construction pipelines, and scalable synthetic data generation systems. Serve as the technical lead and primary point of reference, ensuring full accountability for all research deliverables. Partner with engineering to guarantee well-managed software delivery and reliability at scale across multiple product lines.
Responsibilities
- Lead the design, build, test, and deployment of end-to-end AI solutions for complex document understanding tasks in the legal domain
- Direct the execution of large-scale projects including advanced semantic chunking models for lengthy, non-uniformly structured legal documents with adjustable granularity
- Direct the execution of document enrichment systems with legal and customer-defined taxonomies
- Direct the execution of LLM-based knowledge graph construction pipelines that extract and link heterogeneous legal knowledge
- Direct the execution of scalable synthetic data generation systems
- Serve as the technical lead and primary point of reference, ensuring full accountability for all research deliverables
- Partner with engineering to guarantee well-managed software delivery and reliability at scale across multiple product lines
- Design comprehensive evaluation strategies for both component-level and end-to-end quality, leveraging expert annotation and synthetic data
- Apply robust training methodologies that balance performance with latency requirements
- Lead knowledge distillation initiatives to compress large models into production-ready SLMs
- Maintain scientific and technical expertise through product deliverables, published research, and intellectual property contributions
- Inform Labs shared capabilities and research themes through novel approaches to challenging business problems
- Independently determine appropriate architectures for complex document understanding challenges, balancing accuracy, efficiency, and scalability
- Make critical technical decisions on semantic chunking strategies, document classification approaches, LLM-based knowledge extraction methods, and multi-document reasoning architectures
- Provide input to business stakeholders, mid-to-senior level leadership, and Labs leadership on long-term AI strategy
- Develop in-depth knowledge of TR customers and data infrastructure across multiple products to shape technical roadmaps
- Partner closely with Engineering and Product teams to translate complex legal document understanding challenges into scalable, production-ready solutions
- Engage stakeholders across multiple product lines to deeply understand use case requirements, shaping objectives that align document understanding capabilities with diverse business needs including next-generation search and deep legal research
- Mentor and coach team members with varied ML/NLP abilities, building technical capability across the organization
Requirements
- PhD in Computer Science, AI, NLP, or a related field, or a Master's degree with equivalent research/industry experience
- Demonstrable hands-on experience building and deploying document understanding systems, information extraction pipelines, or knowledge graph construction using deep learning, LLMs, and NLP methods
- Proven ability to translate complex document understanding problems into innovative AI applications that balance accuracy and efficiency
- Demonstrated ability to provide technical leadership, mentor team members, and influence without formal authority in an applied research setting
- Publications at relevant venues such as ACL, EMNLP, ICLR, NeurIPS, SIGIR, or KDD
- Deep understanding of document understanding fundamentals: document layout analysis, semantic chunking approaches beyond fixed-size or paragraph-based methods, document classification handling hierarchical taxonomies, imbalanced multi-label classification, and adapting to domain-specific schemas
- Expertise in knowledge extraction and knowledge graph construction: entity recognition and linking, relation extraction, citation parsing, and building graph representations from unstructured text
- Expertise in LLM-based information extraction, few-shot and multi-task learning, post-training, and knowledge distillation
- Solid understanding of synthetic data generation techniques for NLP, including query-answer generation with verification and scalable data augmentation for training specialized models
- Solid understanding of efficiency optimization including knowledge distillation, model compression, and designing SLM-based solutions that balance performance with computational constraints
- Solid understanding of DL/ML approaches used for NLP tasks
- Experience designing annotation workflows, creating high-quality labeled datasets with clear guidelines, and developing evaluation frameworks for document understanding tasks
Skills
- Python
- PyTorch
- Hugging Face Transformers
- DeepSpeed
- NLP
- GenAI
- Deep Learning
- LLMs
- Document Understanding
- Information Extraction
- Knowledge Graph Construction
- Semantic Chunking
- Document Enrichment
- Knowledge Distillation
- Synthetic Data Generation
- ML
Location
- Zug, Switzerland
- London, UK
Work Type
- Hybrid
Experience Level
- Lead
Education Level
- PhD in Computer Science, AI, NLP, or a related field
- Master's degree with equivalent research/industry experience
Benefits
- Flexible vacation
- Two company-wide Mental Health Days off
- Access to the Headspace app
- Retirement savings
- Tuition reimbursement
- Employee incentive programs
- Resources for mental, physical, and financial wellbeing
- Work from anywhere for up to 8 weeks per year
About the Company
- Thomson Reuters informs the way forward by bringing together the trusted content and technology that people and organizations need to make the right decisions.
- We serve professionals across legal, tax, accounting, compliance, government, and media.
- Our products combine highly specialized software and insights to empower professionals with the data, intelligence, and solutions needed to make informed decisions, and to help institutions in their pursuit of justice, truth, and transparency.
- Reuters, part of Thomson Reuters, is a world leading provider of trusted journalism and news.
- We are powered by the talents of 26,000 employees across more than 70 countries, where everyone has a chance to contribute and grow professionally in flexible work environments.
- At a time when objectivity, accuracy, fairness, and transparency are under attack, we consider it our duty to pursue them.
- Join us and help shape the industries that move society forward.
Equal Opportunity
- As a global business, we rely on the unique backgrounds, perspectives, and experiences of all employees to deliver on our business goals.
- To ensure we can do that, we seek talented, qualified employees in all our operations around the world regardless of race, color, sex/gender, including pregnancy, gender identity and expression, national origin, religion, sexual orientation, disability, age, marital status, citizen status, veteran status, or any other protected classification under applicable law.
- Thomson Reuters is proud to be an Equal Employment Opportunity Employer providing a drug-free workplace.
- We also make reasonable accommodations for qualified individuals with disabilities and for sincerely held religious beliefs in accordance with applicable law.
