We are excited to have a presentation by Lukas Helff from TU Darmstadt on the evolving frontier of reasoning in large language models (LLMs).
Abstract: Reasoning has become the new frontier of artificial intelligence. Recent “reasoning models,” autoregressive LLMs that scale test-time computation, promise more deliberate, human-like “thinking” mechanisms. Yet as these models become more capable, a key question emerges: what does it actually mean for a model to reason? We begin by unpacking this new paradigm of reasoning models, where performance no longer scales with parameter size alone but with how much thinking a model does, that is, how much computation it is provided at test time. A key enabler of such systems is the availability of structured reasoning data and verifiable feedback. This idea is exemplified by “Scalable Logical Reasoning”, a framework that generates symbolic, self-evaluating reasoning tasks with automatic rewards, demonstrating how verifiable structure can make reasoning measurable and trainable.But does scaling test-time compute alone truly lead to genuine reasoning? To explore this, we revisit the roots of AI – Symbolic AI, often referred to as “Good Old-Fashioned AI (GOFAI)” – where reasoning has a long and rigorous history, from symbolic inference and logical reasoning to inductive logic programming. These traditions emphasize structure, explanation, and verifiability, qualities that large language models still lack. Drawing on these roots, we ask whether logical reasoning can not only happen on top of modern language models, but within them. This idea brings us to “Activation Reasoning”, where logical reasoning is embedded directly into the model’s latent activations. By mapping internal representations to logical propositions, AR turns hidden activations into interpretable building blocks for structured and transparent reasoning. In doing so, it bridges classical symbolic ideas of reasoning with the neural foundations of today’s LLMs. Finally, we connect reasoning to AI safety, arguing that the ability to reason and explain is also the foundation of trustworthiness. In multimodal settings, models like LlavaGuard demonstrate how structured reasoning principles can safeguard content by applying flexible policies and providing explicit rationales. This makes reasoning not just a path to better performance, but a foundation for trust. Taken together, we follow reasoning across multiple axes: from scaling reasoning externally through compute, to embedding reasoning internally within activations, and ultimately applying reasoning ethically through safety frameworks. They suggest that the future of AI lies not just in models that think more, but in models that understand why they think.
Speaker Bio: Lukas Helff is a PhD student of Machine Learning at the Computer Science Department of the TU Darmstadt University, Germany.
