EPISODE · May 11, 2026 · 3 MIN
Sakana AI and NVIDIA Introduce TwELL with CUDA Kernels for 20.5% Inference and 21.9% Training Speedup in — 2026-05-11
from Impact Vector: AI Tools · host Alutus LLC
## Short Segments Memori Labs introduces a new way to build persistent memory for AI agents, enhancing multi-user and multi-session applications. Today, we're diving into how Memori's agent-native memory infrastructure allows AI applications to retain context across interactions, making them more effective in real-world scenarios. Later, we'll explore how Sakana AI and NVIDIA's TwELL technology is revolutionizing large language model efficiency. Memori Labs has unveiled a coding implementation that enables AI agents to maintain persistent memory across multiple users and sessions. This development is crucial for building more context-aware applications, as it allows AI models to remember past interactions and user preferences. By integrating Memori into a Google Colab environment, developers can connect it to OpenAI clients, ensuring that every model call passes through this memory layer. The tutorial demonstrates how user data is stored and retrieved, showcasing practical examples like customer-support workflows. This approach helps AI agents retain useful context, moving beyond treating each conversation in isolation. As AI applications become more complex, the ability to maintain context across interactions is increasingly important for delivering personalized and efficient user experiences. ## Feature Story Sakana AI and NVIDIA have introduced TwELL, a breakthrough in large language model efficiency, offering significant speedups in both inference and training. TwELL, which stands for Tile-wise ELLPACK, is an open-source sparse data format and set of CUDA kernels designed to enhance GPU efficiency by skipping ineffective computations. This innovation targets the feedforward layers of large models, where over 80% of neurons remain inactive during text generation. By optimizing GPU operations, TwELL increases inference speed by up to 30% and training speed by up to 24% on H100 GPUs, without compromising model accuracy. The key to TwELL's success lies in its ability to address the inefficiency in feedforward network layers, which account for a significant portion of model parameters and FLOPs. Traditional sparse formats often fail to deliver actual speedups due to the overhead of converting activations from dense to sparse representation. However, TwELL's approach leverages the parallel logic of GPUs, allowing data to be processed in small tiles that GPUs handle efficiently. This method eliminates the need for time-consuming global memory reads and writes, seamlessly integrating into modern chip acceleration pipelines. TwELL's development represents a significant advancement in the field of AI, as it addresses a fundamental bottleneck in scaling large language models. By making computations inside feedforward layers significantly cheaper, TwELL reduces the cost of training and deploying billion-parameter models. This innovation is particularly relevant as the demand for more powerful and efficient AI models continues to grow. As AI researchers and developers seek to push the boundaries of what's possible with large language models, TwELL offers a practical solution to one of the most challenging aspects of model scaling. Looking ahead, the adoption of TwELL could lead to more widespread use of large language models in various applications, from natural language processing to complex decision-making systems. As the AI community continues to explore new ways to optimize model performance, TwELL stands out as a promising development that could reshape the landscape of AI research and deployment. Stay tuned as we follow the impact of TwELL and other innovations in the AI space.
Embed this episode
NOW PLAYING
Sakana AI and NVIDIA Introduce TwELL with CUDA Kernels for 20.5% Inference and 21.9% Training Speedup in — 2026-05-11
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.