Entertainment Providers๐ฅ 10K+ employees employees๐ Los Angeles, California
Job Overview
Join TikTok's Compute Platform SRE team to support all Big Data services and products across the company. This is a newly established team seeking talented individuals to shape its future.
Key Responsibilities
Ensure reliability of TikTok's major data warehouse products, services, and query engines like ClickHouse, Spark, Presto, and Doris
Uphold Service Level Agreements (SLAs) and respond to system outages
Analyze service performance and implement proactive measures to prevent disruptions
Lead incident management and postmortems
Automate infrastructure provisioning and management
Collaborate with product and development teams
Assess and forecast infrastructure needs
Stay updated with industry trends in site reliability engineering
Required Qualifications
Bachelor's Degree or above in Computer Science, Engineering, or related field
Deep understanding of Linux, computer networking, and databases
Proficiency with SRE/DevOps toolsets and Kubernetes
Experience with technologies like ClickHouse, Spark, Doris, Presto, and Kubernetes
Strong coding skills in Python, Shell, Java, or Go
Excellent problem-solving and critical thinking abilities
About TikTok
TikTok is the leading destination for short-form mobile video, inspiring creativity and bringing joy to millions. With global headquarters in Los Angeles and Singapore, TikTok has offices worldwide and a mission to connect people through authentic expression.
Benefits
Employees receive comprehensive benefits including medical, dental, and vision insurance, 401(k) with company match, paid parental leave, disability coverage, life insurance, and wellbeing benefits. The company also provides 10 paid holidays, 10 sick days, and 17 days of Paid Personal Time.