Join us for an in-depth exploration of Large Language Model (LLM) inference, where we'll dive into the intricacies of the computational processes involved in prefill and decode phases, identify algorithmic accelerations such as KVCache and PagedAttention, and explore speedup implementations using CUDA, Triton, and CuTe. We'll discuss techniques like speculative decoding, quantization strategies including NVFP4, and how to effectively use profilers and CUDA Graphs during LLM inference. Depending on the presenter's availability, the event may feature a hands-on session and insights on contributing to vLLM. Additional resources will be provided for those interested in reducing memory usage or exploring multi-node distribution methods. This event is ideal for participants who understand AI classification models and LLM language generation and have a basic grasp of linear algebra. Expect a dynamic and engaging session lasting 3-4 hours. Don't miss this chance to elevate your understanding of LLM inference and network with fellow enthusiasts. RSVP now to secure your spot!