Date/time: 16:00 JST, August 27, 2026
Online Venue: Open to all registered participants. The seminar will be delivered via Zoom. The URL will be provided only to registered participants.
On-site Venue: For RIKEN members only. Open Space at the RIKEN Nihonbashi Office.
Title: Easy-to-Hard Generalization and Small Batch Training for LLMs
Abstract:
In this talk, we will explore both generalization and optimization for neural networks. First, we will discuss work on how to build neural networks that can generalize to far more complex problems than they were trained on by “thinking” for longer. Then, we will discuss recent work on stabilizing optimization for language models. We show that vanilla SGD without momentum is virtually as fast as Adam in the small-batch regime of LLM pretraining.
Bio:
https://goldblum.github.io
Micah Goldblum is an assistant professor in the department of electrical engineering at Columbia University. Before his current position, Micah was a postdoctoral researcher at New York University with Yann LeCun and Andrew Gordon Wilson. Micah’s research focuses on both applied and fundamental problems in machine learning including training, architecture, and inference strategies for large-scale models, AI safety, agents, and building a mathematical and also scientific understanding of why complex AI systems work.
Public events of RIKEN Center for Advanced Intelligence Project (AIP)
Join community