Owning our technology stack end-to-end enables us to optimize every layer - something that becomes increasingly important as we deploy ever-hungrier AI workloads. When we decided to double down on our own data centers and infrastructure rather than continue… more
Owning our technology stack end-to-end enables us to optimize every layer - something that becomes increasingly important as we deploy ever-hungrier AI workloads. When we decided to double down on our own data centers and infrastructure rather than continue migrating to Azure, it really came down to timing and scale. Azure was growing at an incredible rate (and still is!) because customer demand was going through the roof, and LinkedIn was seeing its own skyrocketing growth at the same time. Building fresh for the platform we needed was the right strategic bet for LinkedIn's future. While the internet's ChatGPT moment hadn't happened yet, we already saw companies like Meta reaping the benefits of developing larger and larger AI models, and the leverage from the systems they were building to support them. We bet on building that same kind of foundation ourselves and years later, the payoff is that our engineers can design entire systems end-to-end: intelligently splitting inference workloads across GPUs and CPUs, distilling large models into smaller ones that are very good at solving a specific optimization problem, and going as deep as writing our own GPU kernels with Liger Kernel to squeeze more out of every training run. Owning all of this made our teams more disciplined and more creative at the same time.
Thanks to Frederic Lardinois for a great conversation at the WeAreDevelopers Conference - video to come!like 73celebrate 7love 6insightful 2
As AI moves from experimentation to production, discipline is going to matter just as much as ambition.
At LinkedIn we saw a while ago that the cost trends were clear, and that the future of our products would be powered by increasingly compute-expensive AI… more
As AI moves from experimentation to production, discipline is going to matter just as much as ambition.
At LinkedIn we saw a while ago that the cost trends were clear, and that the future of our products would be powered by increasingly compute-expensive AI workloads that are more intelligent, more capable, and more useful. But those workloads have to be ROI positive, and that’s not something you can just wake up and do overnight. So we set ourselves a challenge that sounds counterintuitive in today’s environment: how do we deliver increasingly sophisticated AI experiences while keeping our compute footprint as close to flat as possible?
This is where owning our own data centers and infrastructure as an applied ML company really pays off, because we can put in the instrumentation and make efficiency a long term investment up and down the stack rather than a one-time push. None of it came from a single breakthrough - it’s the outcome of hundreds of improvements, from optimizing GPU utilization and workload allocation to distilling larger models into smaller and more efficient ones to rethinking how work is distributed across training, inference, storage, and systems design. The results of that intentional push are starting to play out in terms of better experiences for our members and customers while giving our engineers the flexibility to keep innovating.
Raghu and I recently sat down with Paresh Dave at WIRED to talk through all of this, check it out!
https://lnkd.in/guYj674dlike 926insightful 31celebrate 29love 15support 6
How does growing up on a small family farm shape the way you approach technology, healthcare, and public service? Senator Ben Ray Luján joined us on the latest episode of The Messy Middle to discuss leadership, community, and the policies that connect people in an increasingly digital world.
What do Beyoncé, Doja Cat, and former Congressman Patrick McHenry all have in common? 👔
Episode 5 of The Messy Middle out now: https://lnkd.in/e_QHEW5j