
// Open Role at Krafton
Krafton seeks a Senior MLOps Engineer with 8+ years of experience to design, build, and operate large-scale GPU clusters and ML platforms, leveraging advanced AI/ML technologies to enhance game development and service efficiency.
우리는 게이머의 로망을 실현하기 위해, 누구도 가지 않는 길을 갑니다.
예상을 뛰어넘는 과감한 상상력과 기술로, 전 세계 팬들이 잊지 못할 세상을 만들기 위해 담대하게 도전하고 개척합니다.
We pioneer the path to players' dreams.
With bold imagination and breakthrough technology, we create unforgettable worlds for fans across the globe.
[AI Research 본부 비전]
크래프톤 AI Research 본부는 자체 딥러닝 연구를 기반으로 AI 기술의 새로운 가능성을 탐구하고, 게임과 다양한 서비스에 적용 가능한 핵심 AI 기술을 개발합니다. 생성형 AI, 멀티모달 AI, 대규모 언어 모델(LLM) 등 최신 AI 기술을 연구하며, 이를 통해 게임 제작과 사용자 경험을 혁신하는 것을 목표로 합니다.
[Culture Fit]
AI Research 본부는 다양한 배경을 가진 구성원들이 함께 일하며, 수평적이고 활발한 커뮤니케이션 속에서 문제를 해결합니다.
직급과 연차를 넘어 자유롭게 의견을 제시할 수 있으며, 여러 직군과의 협업을 통해 연구와 플랫폼, 서비스의 접점을 함께 만들어갑니다.
팀 소개
KRAFTON MLSys & Ops 팀은 AI Research 본부 내 모델 개발을 위한 GPU 인프라를 설계·구축·운영합니다.
모델 학습 및 실험 환경, GPU 클러스터 운영, 인프라 자동화와 관측성 체계를 함께 다루며, AI 워크로드가 안정적이고 효율적으로 동작할 수 있도록 공통 플랫폼을 만들어갑니다.
이번 포지션은 이미 확보된 B300 125노드 기반 GPU Infrastructure를 함께 운영·고도화하면서, 이를 연구·개발 조직이 안정적이고 효율적으로 활용할 수 있는 GPU Platform으로 발전시키는 실무형 시니어 엔지니어 역할입니다.
완벽하게 모든 경험을 갖춘 분보다는, AI/ML 인프라의 특정 영역을 깊이 있게 다루며 확실한 전문성과 문제 해결 경험을 쌓아온 분을 찾습니다.
ML/GPU Platform 운영 체계 개선 경험
시스템 전반 관점의 문제 해결 역량
요구사항 정의 및 기술적 의사결정 역량
AI 도구를 활용한 업무 생산성 개선 경험
해외 출장에 결격 사유가 없는 분
필수는 아니지만, 아래 영역 중 하나라도 깊이 있게 파고들며 전문성과 성과를 쌓아온 경험이 있다면 팀에 큰 임팩트를 줄 수 있습니다.
차세대 GPU 아키텍처 구축·운영 및 최적화 경험
고성능 GPU 네트워크 분석 및 최적화 경험
분산 스토리지 구축·운영 및 성능 최적화 경험
GPU 자원 활용 및 비용·성능 최적화 경험
GPU 리소스 관리 및 오케스트레이션 도구 활용 경험
분산 학습 성능 개선 경험
- Experience designing, building, and operating large-scale GPU clusters or Kubernetes-based ML platforms for AI/ML workloads. - Proven ability to improve scheduling, multi-tenancy, workload isolation, quotas, observability, and disaster recovery in Kubernetes-based ML/GPU platforms. - Experience implementing cost/performance optimization strategies for GPU utilization, resource allocation, and scheduling. - Deep understanding of the ML workflow (training, experimentation, deployment, serving) and experience operating/advancing common ML platforms. - Track record of establishing repeatable platform operation standards using IaC, GitOps, CI/CD, and observability systems. - Ability to analyze and resolve root causes of issues and implement structural improvements across systems. - Experience collaborating with research, development, and service organizations to define requirements and propose technical solutions for ML/GPU platforms. - Proficiency in using generative AI, LLM-based tools, or code assistants to enhance operational efficiency and productivity. - No disqualification for overseas travel.