📋 Job Overview
TwelveLabs is building AI infrastructure that enables machines to understand video. We are creating more sophisticated and in-depth video understanding AI models that comprehensively analyze visual and auditory information in videos, applying them to workflows in media, entertainment, sports, security, and public sectors. We have raised over 30 billion KRW from global investors including NEA, Radical Ventures, Amazon, NVIDIA, Snowflake, Databricks, Index Ventures, Naver Ventures, Korea Investment Partners, Quadrille Capital, and Red Bull Ventures. We operate offices in San Francisco, Seoul, New York, and London, with core R&D taking place in Seoul.
🏢 About Twelve Labs
TwelveLabs is building AI infrastructure that enables machines to understand video. Our AI models understand video more intricately and deeply than humans, comprehensively analyzing visual and auditory information in videos to apply to workflows in media, entertainment, sports, security, and public sectors. We have raised over 30 billion KRW from global investors including NEA, Radical Ventures, Amazon, NVIDIA, Snowflake, Databricks, Index Ventures, Naver Ventures, Korea Investment Partners, Quadrille Capital, and Red Bull Ventures. We operate offices in San Francisco, Seoul, New York, and London, with core R&D taking place in Seoul.
🎯 The Role
TwelveLabs is looking for a Staff ML Engineer to lead the technical direction and production operations of the Video Ingestion & Serving Platform. This role involves ensuring stable operation of large-scale video processing and GPU workloads, continuously improving throughput, latency, and cost efficiency. You will also design and execute major architecture and migration plans, including data layers and infrastructure.
✅ Key Responsibilities
- Design, develop, and operate backend services and platforms for large-scale video processing, embedding, and AI model serving
- Improve concurrency, retry, backpressure, and fault recovery mechanisms for long-running tasks based on Durable Workflow like Temporal
- Optimize throughput, latency, GPU utilization, and infrastructure costs through Load Testing, Profiling, and operational metrics
- Design PostgreSQL-based data models, Sharding, and Query Paths, leading seamless Production operations
- Build infrastructure and CI/CD with Kubernetes, Terraform, ArgoCD, and improve scalability with Karpenter, KEDA
- Enhance Observability and Incident Response systems including Metrics, Traces, Logs, and Alerting for service Reliability
- Collaborate with Backend, ML, Infrastructure, and Product teams to determine technical direction and lead Production Rollout of major projects
📌 Required Qualifications
- 5+ years of software engineering experience or equivalent capabilities
- Experience designing and operating Distributed Systems including Workflow, Queue, and Database in Production environments
- Proficiency in either Go or Python, with practical development capability in other languages
- Experience operating services on Kubernetes and Cloud Infrastructure, with automation using IaC tools like Terraform
- Experience performing Live Migration of Database, Storage, or Backend systems considering service interruptions and data consistency
- Experience identifying performance bottlenecks and improving throughput or cost through Load Testing, Profiling, and Monitoring
- Experience leading technical decision-making across multiple teams, and improving team execution through design reviews and mentoring
- Able to take high-level Ownership in rapidly changing environments and lead from problem definition to Production operations
⭐ Desirable Experience
- Experience building or operating GPU-based Inference Serving systems like KServe, vLLM, Triton
- Experience with Durable Workflow Engines like Temporal, Cadence, Step Functions
- Experience with Sharded or Distributed Databases like Aurora Limitless, Citus, Vitess
- Experience building and operating Observability Stack based on Grafana, Mimir, Loki, Alloy, OpenTelemetry
- Experience with FFmpeg, video decoding, transcoding, or large-scale media processing pipelines
- Strong communication skills to collaborate with global teams in English
🎁 Benefits
- Global team working with B2B customers
- Hybrid work model combining autonomy and collaboration
- Latest MacBook and equipment worth over 700,000 KRW for remote work, replaced every 3 years
- Unlimited LLM token support for tech teams to use AI efficiently and responsibly
- Annual self-development fund of 1.4 million KRW for lectures, conferences, memberships
- English education programs and global buddy programs
- Evening and weekend taxi fare support
- Corporate card worth 7.2 million KRW annually for meals and transportation
- Snack bar operation in the office (snacks, coffee, seasonal fruits)
- Dinner provided after 7 PM for office-based work
- Annual health check-up for yourself and one family member
- Group insurance enrollment (accident insurance, dental insurance, family accident insurance - choose one)
- Influenza vaccination fee support
- Two-week paid Holiday Break at year-end
🛂 Visa & Eligibility
Common qualification for all positions in Korea: No disqualifying reasons for overseas travel.
All new hires are subject to a 3-month probation period with 100% salary payment during this period.
Candidacy may be canceled or future hiring restricted if false information or misconduct is found in submitted documents.