We’re building toward a world where every company can become its own AI lab. Goaly is a stealth AI startup founded by ex-Meta MSL engineers and researchers. Our mission is to dramatically lower the cost, time, and talent barriers to building proprietary AI — and make each generation of models faster and cheaper to build than the last. Backed by leading AI investors and endorsed by frontier AI researchers and builders, we’re looking for exceptional new grads who want to work on hard, foundational AI systems problems with outsized ownership from day one.
🏢 About Goaly
Goaly is a stealth AI startup founded by ex-Meta MSL engineers and researchers. Our mission is to dramatically lower the cost, time, and talent barriers to building proprietary AI — and make each generation of models faster and cheaper to build than the last. Backed by leading AI investors and endorsed by frontier AI researchers and builders, we’re looking for exceptional new grads who want to work on hard, foundational AI systems problems with outsized ownership from day one.
🎯 The Role
Running modern AI workloads at scale creates systems problems that rarely fit within a single layer of the stack. A slowdown that appears in a training job may originate in a GPU kernel, collective communication, data movement, container runtime, storage path, scheduler, or interaction between model architecture and hardware topology. You will identify these bottlenecks and build systems that improve the throughput, efficiency, and robustness of our largest distributed workloads. Your scope will span post-training, agentic reinforcement learning, model training, rollout inference, and the GPU cluster platform beneath them. You will work closely with researchers and systems engineers, develop a quantitative understanding of performance, and turn one-off investigations into durable infrastructure improvements.
✅ Key Responsibilities
Profile end-to-end AI workloads and identify limiting resources across model code, GPU kernels, memory, collective communication, networking, storage, orchestration, and environment execution.
Build low-latency, high-throughput sampling and inference systems for large language models, including batching, scheduling, caching, load balancing, and efficient weight updates.
Optimize GPU execution through kernel and graph profiling, memory-layout improvements, reduced-precision computation, communication overlap, compilation, and targeted CUDA or Triton work.
Improve distributed training and reinforcement-learning performance across heterogeneous GPU and CPU workloads, variable-length rollouts, complex network topologies, and changing model architectures.
Design quantitative performance and capacity models that predict bottlenecks, explain scaling behavior, guide hardware and topology choices, and prioritize engineering work.
Build scheduling and load-balancing mechanisms that improve accelerator utilization while respecting memory, locality, topology, latency, and fault-domain constraints.
Design fault-tolerant distributed systems that detect failures early, isolate their impact, recover efficiently, and preserve correctness during long-running jobs.
Investigate difficult production issues such as kernel-level stalls, network-latency spikes, collective timeouts, memory fragmentation, stragglers, and performance regressions in containerized environments.
Develop benchmarks, profiling tools, performance dashboards, and regression tests that make system behavior visible and allow improvements to be measured under realistic workloads.
Partner with researchers to understand new models and algorithms, remove infrastructure constraints from the experimental loop, and translate successful optimizations into reusable platform capabilities.
📌 Required Qualifications
Significant software-engineering, distributed-systems, high-performance-computing, or ML-infrastructure experience, particularly with performance-critical systems operating at large scale.
Exceptional programming and debugging ability in Python and at least one systems language such as C++, Rust, or Go.
Strong systems fundamentals, including operating systems, concurrency, memory, networking, storage, scheduling, containerization, and failure handling.
A track record of using measurement to solve ambiguous performance problems: forming hypotheses, designing representative benchmarks, reading profiles and traces, identifying root causes, and validating improvements under production conditions.
The ability to reason across abstraction boundaries, from model architecture and framework execution to accelerator behavior, distributed runtimes, cluster topology, and infrastructure services.
A results-oriented mindset, flexibility about where in the stack to work, and a willingness to take ownership beyond a narrowly defined job description.
Strong communication and collaboration skills, including the ability to work directly with researchers, explain complex systems behavior clearly, and turn repeated investigations into maintainable tools and abstractions.
Interest in developing deep expertise in machine learning systems, even if your prior work has been primarily in distributed system
⭐ Desirable Experience
Experience with large-scale distributed systems and high-performance computing.
Experience with machine learning frameworks and tools.
Experience with GPU programming and optimization.
Experience with containerization and orchestration tools.
Experience with performance profiling and benchmarking tools.
🎁 Benefits
Competitive salary and benefits package.
🛂 Visa & Eligibility
Eligibility to work in the United States required.