About the Role
LG AI Research is leading the AI transformation across industries through the research and development of large-scale AI models, including EXAONE. The Data Governance team within the Product Unit establishes policies for data generated and utilized across the research organization. We design and build an integrated data platform for consistent collection, refinement, version management, search, sharing, and access control. Based on this platform, we implement Practical AI by applying the latest AI techniques to actual services and tasks across various industries. Our goal is to solve the limitations in search, versioning, sharing, and governance that arise from multiple research labs independently managing data. We aim to create an environment where researchers can focus on their work without data-related concerns by stably operating petabyte-scale research data. We are looking for someone to proactively design and implement this direction from a data platform architecture perspective.
Responsibilities
- Design Data Catalog (search/classification/quality/ownership) and standard schema/policy frameworks.
- Design and implement data access control (permissions/policies/token·IAM-based) and audit/tracking systems.
- Develop SDK/CLI and integrate them into research workflows (training/evaluation/experimentation).
- Deploy and operate Kubernetes-based Data Pipeline services, including CI/CD and operational automation.
- Design and operate PB-scale storage and pipeline architectures based on Google Cloud Platform.
- Optimize performance, cost, and availability for multi-organization/large-scale data, operating based on SLO/SLA.
- Design system monitoring/observability environments (metrics/logs/alerts) and enhance incident response systems.
Requirements
- Over 7 years of software engineering experience, with at least 3 years in data infrastructure/platforms and experience in data platform architecture design.
- Experience designing and operating infrastructure on Google Cloud Platform.
- Experience deploying and operating Kubernetes-based services (including operational automation/incident response).
- Proficiency in Python-based backend/platform development and experience designing/implementing large-scale systems.
- Understanding of distributed system design principles and operational experience.
- Experience designing and developing structured/unstructured data pipelines (collection/processing/extraction).
- Familiarity with development and operations in Linux environments, with the ability to communicate effectively with various stakeholders.
Skills
- Data Catalog
- Data Governance
- Data Platform Architecture
- Google Cloud Platform
- Kubernetes
- Python
- Backend Development
- Platform Development
- Large-scale Systems
- Distributed Systems
- Data Pipelines
- Linux
- SDK
- CLI
- CI/CD
- IaC
- GitOps
- Data Versioning
- Lakehouse
- Apache Spark
- Trino
- Flink
- Helm
- Terraform
- ArgoCD
Location
- South Korea
Work Type
- Full-time
Experience Level
- 7+ years of software engineering experience
- 3+ years in data infrastructure/platforms
About the Company
- LG AI Research is leading the AI transformation across industries through the research and development of large-scale AI models, including EXAONE.
- The Data Governance team within the Product Unit establishes policies for data generated and utilized across the research organization.
- We design and build an integrated data platform for consistent collection, refinement, version management, search, sharing, and access control.
- Based on this platform, we implement Practical AI by applying the latest AI techniques to actual services and tasks across various industries.
- Our goal is to solve the limitations in search, versioning, sharing, and governance that arise from multiple research labs independently managing data.
- We aim to create an environment where researchers can focus on their work without data-related concerns by stably operating petabyte-scale research data.
